MicroGPT

A Language Model Small Enough to Read

4,192 parameters, 199 lines of pure Python — the same structure as a large model, only smaller
Large language model MicroGPT Vocabulary: over 100,000 tokens 27 characters (a–z plus BOS) Vectors: thousands of dimensions 16 dimensions Dozens of layers · dozens of heads 1 layer · 4 heads Task: continue any text Task: learn names, write the next letter
← → to turn pages · f for fullscreen
explanation of the great work karpathy.github.io/2026/02/12/microgpt

THE RESULT FIRSTWhat running this program produces

Reads 32,033 real names, writes 20 that do not exist
$ python microgpt.py num docs: 32033 vocab size: 27 num params: 4192 step 1000 / 1000 | loss 2.6497 --- inference (new, hallucinated names) --- sample 1: kamon sample 2: ann sample 3: karai sample 4: jaire sample 5: vialan sample 6: karia sample 7: yeran sample 8: anna sample 9: areli sample 10: kaina sample 11: konna sample 12: keylen sample 13: liole sample 14: alerin sample 15: earan sample 16: lenne sample 17: kana sample 18: lara sample 19: alela sample 20: anton 32,033 real English names are read in A vocabulary of 27: 26 letters plus BOS The whole model holds only 4,192 numbers Loss: lower means better predictions None of these 20 names exists Not one of the 32,033 names read in looks like these. Every one was written by the model, one letter at a time. What this talk is about What happens inside this program, from reading the names in to writing the 20 names above. This talk covers only the next-letter step. Later pages use emma from the dataset as the running example.

CONCEPT 1Autoregression: one token at a time

The output becomes the input of the next round
This example uses a character-level tokenizer: one letter per token, with BOS marking the start of a sequence. Text so far (context) BOS e m One pass The whole context is read Next token m the second m in emma Appended to the context; the model runs again Round by round One token per round (gold = new in that round) Round 1 BOS e m m Round 2 BOS e m m a Round 3 BOS e m m a BOS BOS appearing again means the name is finished; the loop stops there. Why text is split into tokens at all, and why it is split this way: see Concept 3. In plain terms Before each new letter, the text so far is read again, and only then is the next letter chosen. A whole reply is not formed in one go. It is one process repeated hundreds or thousands of times. In practice the intermediate results (K and V) of each round are kept in a KV cache and reused.

CONCEPT 2The five steps of inference

Text goes in, one token comes out
Every token produced runs through all five steps This page covers a single round; for how the new token feeds back, see Concept 1. ① Tokenize Split the text into numbered tokens ② Embedding Look each number up and get a vector ③ Transformer ×N Attention looks back MLP refines Vectors now carry context ④ Score Give every token a probability ⑤ Sample Draw one token by probability see Concept 3 see Concept 4 see Concepts 5 and 6 see Concept 7 The next five pages zoom in on ① through ⑤ in turn.

CONCEPT 3① Tokenize: text becomes numbers

One fixed list is chosen up front, and nothing outside that list can ever be produced
ⓐ Why this step exists The model works on numbers, not on text. text numbers reading in writing out Both directions use one fixed list: the vocabulary. Once set, the list is fixed Anything outside it cannot be read in, and can never be written out. ⓑ The split used here: characters 27 entries One character is one token. Vocabulary: 26 letters plus BOS a b c … m … z 0 1 2 12 25 BOS 26 BOS marks where a name starts and ends. Example emma BOS e m m a 26 4 12 12 0 From here on the model sees only these numbers. Every later page uses the same moment: emma, about to write the second m. ⓒ The trade-off in the split Fine or coarse: the granularity is a choice. Character level (here) 27 entries e m m a One letter per round, over several rounds. Subword level (large models) 100,000+ emm a A whole word stem in a single round. The five steps stay identical. Only the size and grain of the list change. This choice fixes how much text one token holds.

CONCEPT 4② Embedding: numbers become vectors

Two lookup tables, one lookup each, then a single addition
ⓐ wte: which token this is 27 rows × 16 A table of 27 rows, one row per token, each row holding 16 numbers. number 12 take row 12 …16 in all A lookup really is that literal: no arithmetic, just row 12 copied out. What are those 16 numbers? Nobody ever assigned them a meaning; every one came from training. Characters used in similar ways end up with similar numbers. Inside the model, the meaning of a character is a point in a 16-dimensional space. ⓑ wpe: which position this is 16 rows × 16 emma has two m's, and the 16 numbers wte returns for them are identical. BOS e m m a 0 1 2 3 4 ← position The positions differ, and so do the roles: one follows e, the other follows m. The model must tell them apart to choose the next letter. So a second table is kept: one row per position, again 16 numbers each. position 0 … position 1 … Rows in the table = the context limit 16 rows here, so at most 16 positions. That number is the context length. Large models have vastly more rows. ⓒ Adding the two together wte row 12 + wpe row 2 vector entering ③ Element by element: first plus first, and so on. Worth pausing on The whole of step ② is two table lookups and one addition. Every calculation that follows rests on the vector that addition produces. After the addition, one vector carries both which character it is and where it sits. Only then does it start reading the others.

CONCEPT 5Three things inside a Transformer

Attention shares · MLP refines · the input is added back
ⓐ Attention: drawing on others sideways The vector from Concept 4 has read no context. BOS e m information from earlier positions The vector for m asks every earlier position for information, takes more from the more relevant ones, and combines them into one vector. BOS and e only supply; they are not recomputed. This is the only place positions exchange anything. How that is actually computed — three roles, similarity, several heads — see Concept 6. ⓑ MLP: each position alone 16 → 64 → 16 After attention, every position passes through the same small network, with no exchange. 16 numbers expanded into 64 checks ReLU back down to 16 Expanding runs 64 checks at once — say, whether the previous letter is a vowel, or whether the end is near. ReLU keeps the positive results and zeroes the negative ones. One across, one down Attention: positions share sideways. MLP: each position works downward. A Transformer block is these two stages. ⓒ Adding back to the original residual · ×N original keep a copy RMSNorm Attention or MLP + the new, combined vector This runs twice inside a single block: once for attention, once for the MLP. Added, not replaced The result is added back on top, so nothing is lost; shapes in and out match, which is why the block can repeat N times. RMSNorm Pulls values back to a comparable scale.

CONCEPT 6Attention up close

Three roles · how much to take · several heads
① How the three roles arise Wq · Wk · Wv the vector for m each multiplied by its own table: Q what is being looked for Wq K what this position is Wk V what it can supply Wv Wq, Wk and Wv are three weight tables. Training fills them in; nobody writes them by hand. Every earlier position likewise forms its own Q, K and V. K and V computed in earlier rounds are kept and reused (the KV cache, Concept 1). ② How much to take from each similarity The Q of m is compared with the K of every earlier position. Closer directions score higher. close direction → high score diverging → low score Each position gets a score, and the scores are then turned into weights that add up to 1. A high score means more of that position's V. The weighted Vs are added into the new vector. ③ Several heads in parallel 4 heads · wo one vector head 1 head 2 head 3 head 4 each head runs ① and ② on its own four results joined into one then through the wo table Heads attend to different things: some to neighbouring letters, some further back. That division of labour emerges from training; no head is assigned a job. wo is likewise a weight table learned in training.

CONCEPT 7④ Scoring and ⑤ Sampling

From one vector to a letter on the page
ⓐ What comes out is one vector Only one vector goes through the Transformer this round: the one for m. vector for m Transformer ×N K, V of BOS and e …16 in all Out come 16 numbers again, but this time they have read BOS and e. BOS and e were computed in earlier rounds. Their K and V sit in the KV cache and are used directly, with no second pass through the Transformer. Each round yields exactly one token, appended after the last letter so far. ⓑ Vector to probability lm_head · softmax …16 in all lm_head (27 × 16) m 27 characters, one probability each, adding up to 1 The vector passes through lm_head and comes out as 27 scores, one per character, which softmax turns into probabilities. Step ② in reverse wte turns a character into a vector; lm_head turns a vector back into scores. ⓒ Sampling the only randomness One draw from these probabilities picks a token. whichever band the draw lands in is the character m this round: the second m in emma Every step before this is fixed arithmetic: the same input always gives the same scores. Temperature One number set before sampling: the 27 scores are divided by it first. Low → sharper and safer; high → flatter. Drawing BOS means this name is finished. Otherwise the new token joins the context and step ① starts another round (Concept 1).

CONCEPT 8Where the 4,192 parameters are

A model holds two kinds of thing: stored numbers, and fixed operations
Parameters: the numbers actually stored Names match the variables in the MicroGPT source. numbers wte character table: 27 characters × 16 432 wpe position table: 16 positions × 16 256 Wq Wk Wv the three tables that produce Q, K and V 768 wo the table applied once the heads are joined 256 MLP expand 16→64, then squeeze 64→16 2048 lm_head 27 character scores computed from a vector 432 Total 4192 No parameters: fixed operations RMSNorm pulls a vector's values to a comparable scale ReLU keeps positives, zeroes negatives softmax turns a set of scores into probabilities + adds two vectors element by element, merging them sampling draws one token by probability These steps store no numbers. They follow the same rules every time, in any model. Next: the same pipeline, walked through once more in 3D The 3D walkthrough follows the first m in emma — the same vector traced on these pages — stage by stage through the five steps.