MicroGPT
A Language Model Small Enough to Read
4,192 parameters, 199 lines of pure Python — the same structure as a large model, only smaller
Large language model
MicroGPT
Vocabulary: over 100,000 tokens
27 characters (a–z plus BOS)
Vectors: thousands of dimensions
16 dimensions
Dozens of layers · dozens of heads
1 layer · 4 heads
Task: continue any text
Task: learn names, write the next letter
←
→
to turn pages ·
f
for fullscreen
explanation of the great work
karpathy.github.io/2026/02/12/microgpt
THE RESULT FIRST
What running this program produces
Reads 32,033 real names, writes 20 that do not exist
$ python microgpt.py
num docs: 32033
vocab size: 27
num params: 4192
step 1000 / 1000 | loss 2.6497
--- inference (new, hallucinated names) ---
sample 1: kamon
sample 2: ann
sample 3: karai
sample 4: jaire
sample 5: vialan
sample 6: karia
sample 7: yeran
sample 8: anna
sample 9: areli
sample 10: kaina
sample 11: konna
sample 12: keylen
sample 13: liole
sample 14: alerin
sample 15: earan
sample 16: lenne
sample 17: kana
sample 18: lara
sample 19: alela
sample 20: anton
32,033 real English names are read in
A vocabulary of 27: 26 letters plus BOS
The whole model holds only 4,192 numbers
Loss: lower means better predictions
None of these 20 names exists
Not one of the 32,033 names read in
looks like these. Every one was written
by the model, one letter at a time.
What this talk is about
What happens inside this program,
from reading the names in to writing
the 20 names above.
This talk covers only the next-letter step.
Later pages use emma from the dataset as the running example.
CONCEPT 1
Autoregression: one token at a time
The output becomes the input of the next round
This example uses a character-level tokenizer: one letter per token, with BOS marking the start of a sequence.
Text so far (context)
BOS e m
One pass
The whole context is read
Next token
m
the second m in emma
Appended to the context; the model runs again
Round by round
One token per round (gold = new in that round)
Round 1
BOS e m
m
Round 2
BOS e m m
a
Round 3
BOS e m m a
BOS
BOS appearing again means the name is finished; the loop stops there.
Why text is split into tokens at all, and why it is split this way: see Concept 3.
In plain terms
Before each new letter, the text so far is read
again, and only then is the next letter chosen.
A whole reply is not formed in one go. It is one
process repeated hundreds or thousands of times.
In practice the intermediate results (K and V) of
each round are kept in a KV cache and reused.
CONCEPT 2
The five steps of inference
Text goes in, one token comes out
Every token produced runs through all five steps
This page covers a single round; for how the new token feeds back, see Concept 1.
① Tokenize
Split the text into
numbered tokens
② Embedding
Look each number
up and get a vector
③ Transformer ×N
Attention looks back
MLP refines
Vectors now carry context
④ Score
Give every token
a probability
⑤ Sample
Draw one token
by probability
see Concept 3
see Concept 4
see Concepts 5 and 6
see Concept 7
The next five pages zoom in on ① through ⑤ in turn.
CONCEPT 3
① Tokenize: text becomes numbers
One fixed list is chosen up front, and nothing outside that list can ever be produced
ⓐ Why this step exists
The model works on numbers, not on text.
text
numbers
reading in
writing out
Both directions use one fixed list: the vocabulary.
Once set, the list is fixed
Anything outside it cannot be read
in, and can never be written out.
ⓑ The split used here: characters
27 entries
One character is one token.
Vocabulary: 26 letters plus BOS
a
b
c
…
m
…
z
0
1
2
12
25
BOS
26
BOS marks where a name starts and ends.
Example
emma
BOS
e
m
m
a
26 4 12 12 0
From here on the model sees only these numbers.
Every later page uses the same moment:
emma, about to write the second m.
ⓒ The trade-off in the split
Fine or coarse: the granularity is a choice.
Character level (here)
27 entries
e
m
m
a
One letter per round, over several rounds.
Subword level (large models)
100,000+
emm
a
A whole word stem in a single round.
The five steps stay identical. Only
the size and grain of the list change.
This choice fixes how much text one token holds.
CONCEPT 4
② Embedding: numbers become vectors
Two lookup tables, one lookup each, then a single addition
ⓐ wte: which token this is
27 rows × 16
A table of 27 rows, one row per token,
each row holding 16 numbers.
number 12
take row 12
…16 in all
A lookup really is that literal: no
arithmetic, just row 12 copied out.
What are those 16 numbers?
Nobody ever assigned them a
meaning; every one came from training.
Characters used in similar ways
end up with similar numbers.
Inside the model, the meaning of a character
is a point in a 16-dimensional space.
ⓑ wpe: which position this is
16 rows × 16
emma has two m's, and the 16 numbers
wte returns for them are identical.
BOS
e
m
m
a
0
1
2
3
4
← position
The positions differ, and so do the roles:
one follows e, the other follows m. The model
must tell them apart to choose the next letter.
So a second table is kept: one row per
position, again 16 numbers each.
position 0
…
position 1
…
Rows in the table = the context limit
16 rows here, so at most 16 positions.
That number is the context length.
Large models have vastly more rows.
ⓒ Adding the two together
wte row 12
+
wpe row 2
vector entering ③
Element by element: first plus first, and so on.
Worth pausing on
The whole of step ② is two table
lookups and one addition.
Every calculation that follows rests
on the vector that addition produces.
After the addition, one vector carries both
which character it is and where it sits.
Only then does it start reading the others.
CONCEPT 5
Three things inside a Transformer
Attention shares · MLP refines · the input is added back
ⓐ Attention: drawing on others
sideways
The vector from Concept 4 has read no context.
BOS
e
m
information from
earlier positions
The vector for m asks every earlier position
for information, takes more from the more
relevant ones, and combines them into one vector.
BOS and e only supply; they are not recomputed.
This is the only place positions exchange anything.
How that is actually computed — three roles,
similarity, several heads — see Concept 6.
ⓑ MLP: each position alone
16 → 64 → 16
After attention, every position passes through
the same small network, with no exchange.
16 numbers
expanded into 64 checks
ReLU
back down to 16
Expanding runs 64 checks at once — say,
whether the previous letter is a vowel, or
whether the end is near. ReLU keeps the
positive results and zeroes the negative ones.
One across, one down
Attention: positions share sideways.
MLP: each position works downward.
A Transformer block is these two stages.
ⓒ Adding back to the original
residual · ×N
original
keep a copy
RMSNorm
Attention
or MLP
+
the new, combined vector
This runs twice inside a single block:
once for attention, once for the MLP.
Added, not replaced
The result is added back on top, so
nothing is lost; shapes in and out match,
which is why the block can repeat N times.
RMSNorm
Pulls values back to a comparable scale.
CONCEPT 6
Attention up close
Three roles · how much to take · several heads
① How the three roles arise
Wq · Wk · Wv
the vector for m
each multiplied by its own table:
Q
what is being looked for
Wq
K
what this position is
Wk
V
what it can supply
Wv
Wq, Wk and Wv are three weight tables.
Training fills them in; nobody writes them by hand.
Every earlier position likewise forms
its own Q, K and V.
K and V computed in earlier rounds are
kept and reused (the KV cache, Concept 1).
② How much to take from each
similarity
The Q of m is compared with the K of every
earlier position. Closer directions score higher.
close direction → high score
diverging → low score
Each position gets a score, and the scores are
then turned into weights that add up to 1.
A high score means more of that position's V.
The weighted Vs are added into the new vector.
③ Several heads in parallel
4 heads · wo
one vector
head 1
head 2
head 3
head 4
each head runs ① and ② on its own
four results joined into one
then through the wo table
Heads attend to different things: some to
neighbouring letters, some further back.
That division of labour emerges from
training; no head is assigned a job.
wo is likewise a weight table learned in training.
CONCEPT 7
④ Scoring and ⑤ Sampling
From one vector to a letter on the page
ⓐ What comes out is one vector
Only one vector goes through the
Transformer this round: the one for m.
vector for m
Transformer ×N
K, V of
BOS and e
…16 in all
Out come 16 numbers again, but this
time they have read BOS and e.
BOS and e were computed in earlier rounds. Their
K and V sit in the KV cache and are used directly,
with no second pass through the Transformer.
Each round yields exactly one token,
appended after the last letter so far.
ⓑ Vector to probability
lm_head · softmax
…16 in all
lm_head (27 × 16)
m
27 characters, one probability each, adding up to 1
The vector passes through lm_head and comes
out as 27 scores, one per character, which
softmax turns into probabilities.
Step ② in reverse
wte turns a character into a vector;
lm_head turns a vector back into scores.
ⓒ Sampling
the only randomness
One draw from these probabilities picks a token.
whichever band the draw lands in is the character
m
this round: the second m in emma
Every step before this is fixed arithmetic: the
same input always gives the same scores.
Temperature
One number set before sampling: the
27 scores are divided by it first.
Low → sharper and safer; high → flatter.
Drawing BOS means this name is finished.
Otherwise the new token joins the context
and step ① starts another round (Concept 1).
CONCEPT 8
Where the 4,192 parameters are
A model holds two kinds of thing: stored numbers, and fixed operations
Parameters: the numbers actually stored
Names match the variables in the MicroGPT source.
numbers
wte
character table: 27 characters × 16
432
wpe
position table: 16 positions × 16
256
Wq Wk Wv
the three tables that produce Q, K and V
768
wo
the table applied once the heads are joined
256
MLP
expand 16→64, then squeeze 64→16
2048
lm_head
27 character scores computed from a vector
432
Total
4192
No parameters: fixed operations
RMSNorm
pulls a vector's values to a comparable scale
ReLU
keeps positives, zeroes negatives
softmax
turns a set of scores into probabilities
+
adds two vectors element by element, merging them
sampling
draws one token by probability
These steps store no numbers. They follow the
same rules every time, in any model.
Next: the same pipeline, walked through once more in 3D
The 3D walkthrough follows the first m in emma — the same vector traced on these pages — stage by stage through the five steps.