Overview
A decoder-only transformer language model built from first principles. The architecture follows GPT-2 Small: 12 transformer layers, 12 attention heads, 768 embedding dimensions, and a 1024-token context window. Trained on 100K+ synthetic customer service conversations using a custom BPE tokenizer.
Model Architecture
Specifications
| Component | Value |
|---|---|
| Vocabulary Size | 5,000 tokens (BPE) |
| Context Length | 1,024 tokens |
| Embedding Dimension | 768 |
| Transformer Layers | 12 |
| Attention Heads | 12 |
| Head Dimension | 64 (768 / 12) |
| Feed-Forward Hidden | 3,072 (4 × 768) |
| Dropout | 0.2 |
| Parameters | ~85 million |
Component Stack
The model is composed of five core modules:
-
EmbeddingModule — Token embeddings + learned positional embeddings, summed to produce input representations of shape
(B, T, 768) -
SelfAttentionHead — Single attention head with Q/K/V projections. Uses a causal mask (lower triangular) to prevent attending to future tokens. Scaled dot-product attention with
head_size^-0.5normalization. -
MultiHeadAttention — 12 parallel attention heads, concatenated and projected back to 768 dimensions
-
FeedForward — Two-layer MLP with ReLU activation:
768 → 3072 → 768with dropout -
TransformerBlock — Pre-LN architecture:
LayerNorm → Attention → Residual → LayerNorm → FFN → Residual
Forward Pass
Input IDs (B, T)
↓
Token Embedding + Positional Embedding → (B, T, 768)
↓
12× TransformerBlock
↓
Final LayerNorm
↓
Linear Head → Logits (B, T, vocab_size)
↓
Cross-Entropy Loss (against shifted targets)
Tokenization
Custom BPE tokenizer trained with Hugging Face tokenizers library:
| Token | ID |
|---|---|
<|unk|> | Unknown token |
<|pad|> | Padding (ignored in loss) |
<|user|> | User turn marker |
<|assistant|> | Assistant turn marker |
<|endoftext|> | End of conversation |
Training data format:
<|user|>
Hi, I just received my latest invoice and there's a charge for "Premium Support" that I don't remember signing up for. Can you help me figure out what's going on?
<|assistant|>
Absolutely, I'd be happy to help. Let me pull up your account details. Could you confirm the invoice number or the date it was sent?
<|user|>
Sure, the invoice number is INV-2025-0789 and it was issued on September 12th.
<|assistant|>
Thanks! I see the invoice in our system. The "Premium Support" line appears because a trial subscription for that service was activated automatically when you upgraded your plan last month. Did you notice any changes in your plan around that time?
<|user|>
I didn't realize there was a trial. I thought the upgrade was just to get more storage.
<|assistant|>
That makes sense. The trial is free for 30 days, but the billing cycle starts after that period. I can cancel the Premium Support trial for you and adjust the invoice so you're not charged. Would that be okay?
<|user|>
Yes, please cancel it. Also, can you confirm that I won't see any additional charges for that service in the next month?
<|assistant|>
I've canceled the trial, and I've updated your invoice to remove the Premium Support line. You won't see any more charges for that service, and I've sent a revised invoice to your email. Let me know if there's anything else I can help with.
<|endoftext|>
Loss masking: User turns are masked with -100 so the model only learns to predict assistant responses, not echo user input.
Training Configuration
| Hyperparameter | Value |
|---|---|
| Batch Size | 32 |
| Learning Rate | 1e-4 |
| Optimizer | AdamW (weight decay: 0.1) |
| LR Schedule | Cosine annealing with 100-step warmup |
| Gradient Accumulation | 4 steps |
| Epochs | 500 (with early stopping) |
| Early Stopping Patience | 5 epochs |
| Eval Interval | Every 100 steps |
Training Loop Highlights
- Gradient accumulation: Effective batch size of 128 (32 × 4) while staying within VRAM limits
- Learning rate warmup: Linear warmup over 100 steps before cosine decay
- Checkpointing: Model, optimizer, and scheduler states saved after each epoch
- Early stopping: Training halts if validation loss doesn't improve for 5 consecutive epochs
Inference & Generation
Sampling Parameters
| Parameter | Default | Description |
|---|---|---|
| max_tokens | 50 | Maximum tokens to generate |
| temperature | 0.7 | Softmax temperature (higher = more random) |
| top_k | 50 | Consider only top-k most likely tokens |
| top_p | 0.9 | Nucleus sampling threshold |
| repetition_penalty | 1.2 | Penalty applied to repeated tokens |
Generation Algorithm
- Encode prompt with
<|user|>...<|assistant|>format - Forward pass → get logits for last position
- Apply repetition penalty to previously generated tokens
- Scale by temperature
- Apply top-p (nucleus) filtering OR top-k filtering
- Sample from resulting distribution
- Append token, repeat until
<|endoftext|>or max tokens
Deployment Stack
| Layer | Technology |
|---|---|
| Inference API | Flask with CORS, streaming SSE |
| Frontend | React + TypeScript |
| Containerization | Docker (CPU and GPU variants) |
| Model Serving | PyTorch with torch.no_grad() context |
Try It Yourself
The demo exposes all generation parameters—temperature, top-k, top-p, and repetition penalty—so you can see how each affects output quality.



