MicroGPT

How 4,192 Numbers Learn to Name People

A 200-line, dependency-free GPT in pure Python — taken apart until only addition and multiplication are left
BOS = Beginning of Sequence BOS Tokenizer Embedding Attention + MLP lm_head Samplingrandomness e The letter drawn becomes the input of the next round (autoregression)
4,192
parameters in total
200
lines of pure Python
1
transformer layer
16
dimensions per vector
27
vocab (a–z + BOS)
← → to turn pages · n speaker notes · f fullscreen · running example throughout: writing emma

OpeningWhat this program actually does

From start to finish the model does one thing letters seen so far e m the model next-letter probabilities 27 of them a41% m22% e9% ⋮ z0.1% It does not decide an answer; it only supplies 27 proportions. One line outside the model draws the letter: random.choices Training and inference both call the same gpt(): most of this talk is one pass with the stored numbers; learning is a loop outside that function that goes back and adjusts those numbers. What the data looks like, and how it turns into numbers — see Fig 1.
microgpt_en.py · the only signature in the talkin → out
108	def gpt(token_id, pos_id, keys, values):
	# ★ in: this letter, and its position
⋯
143	    logits = linear(x, state_dict['lm_head'])
144	    return logits
	# ★ out: 27 raw scores, not yet probabilities
The task, whole talk: text so far in → 27 probabilities for the next letter out (softmax(logits) turns scores into probabilities)

Fig 0How small this model is

Everything the model knows = 5 groups of parameters Each matrix is a table of numbers. Rows × columns = cells; all of them add up to the 4,192 below. wte 27 × 16 · letter → vector (lookup) 432 wpe 16 × 16 · one vector per position, 0–15 256 lm_head 27 × 16 · vector → letter (lookup in reverse) 432 attention Wq Wk Wv Wo · 16 × 16 each 1,024 MLP 16 → 64 → 16 2,048 Total 4,192 numbers — every one countable by hand 432+256+432+1,024+2,048 And only one layer — that layer already contains all of attention, MLP and the residual; more layers are the same thing done again. For comparison GPT-2 small 124,000,000 × 30,000 · not to scale LLMs today 100,000,000,000+
microgpt_en.py · parameter initnum params: 4192
75	n_layer = 1     # depth of the transformer …
76	n_embd = 16     # width of the network …
77	block_size = 16 # maximum context length …
78	n_head = 4      # number of attention heads
79	head_dim = n_embd // n_head # derived dimension …
	# ★ 5 hyper-parameters: 1 layer, 16 wide, 4 heads
80	matrix = lambda nout, nin, std=0.08:
	↪ [[Value(random.gauss(0, std)) for _ in range(nin)]
	↪  for _ in range(nout)]
	# ★ matrix: one table; Value = number + notepad
81	state_dict = {'wte': matrix(vocab_size, n_embd),
	↪ 'wpe': matrix(block_size, n_embd),
	↪ 'lm_head': matrix(vocab_size, n_embd)}
	# ★ state_dict: where all knowledge is kept
82	for i in range(n_layer):
83	    state_dict[f'layer{i}.attn_wq'] =
	↪         matrix(n_embd, n_embd)
⋯
87	    state_dict[f'layer{i}.mlp_fc1'] =
	↪         matrix(4 * n_embd, n_embd)
⋯
89	params = [p for mat in state_dict.values()
	↪ for row in mat for p in row] # flatten …
	# ★ params: all 4,192 numbers in one flat list
90	print(f"num params: {len(params)}")

Warm-up 0Only two kinds of number in the code

① Learned (parameters) ② Computed (intermediate values) wte wpe lm_head Wq Wk Wv Wo fc1 fc2 x q k v logits probs 4,192 of them — the table on the previous page unchanged however many times it runs they live in state_dict only training ever changes them computed from scratch every round recomputed for every name usually discarded once used; k and v are no exception the cache only keeps them alive across rounds what the model knows = its knowledge what the model is working on = its train of thought Any new variable comes down to one question: left side or right side? Anything starting with state_dict[...] is on the left; nearly everything else is on the right. Training tunes the left; inference uses the left to compute the right. The exception is keys / values: they belong on the right, yet are kept across rounds (the KV cache, Fig 9).

Warm-up 1Every call moves along one 16-dimensional x

In form it is a Python list holding 16 decimals: x = [ 0.13, -0.82, 0.55, … , -0.07 ] These 16 numbers condense what is understood about this letter and every letter before it After three additions, before the exit (lm_head): the meaning at this moment and the prediction of the next letter sit in the same 16 numbers — Fig 7 shows the two are one and the same The verb at every step is + — three additions in all The whole program keeps adding to one x (writing emma, now at step 3): born · wte lookup “stands for the letter m” lookup + wpe position known “the m at position 3” lookup + what the earlier text adds earlier text mixed in “preceded by e, m” attention + what refining adds pass complete “a is likely to come next” MLP Every figure that follows answers the same question: “What did this step add to x?” From start to finish, only three operations ever touch this x

Warm-up 2① Addition: two vectors, two ideas stacked

The model uses only three operations. The next three pages take one each, on the same example. “stands for the letter m” [ 0.3, -0.1, 0.7, … ] ← a row of wte + “position 3” [ 0.1, 0.4, -0.2, … ] ← a row of wpe “the m at position 3” [ 0.4, 0.3, 0.5, … ] two notes stacked A second use of addition: laying a correction on top (what attention / MLP produced) x + correction the updated x new x = old x + correction (the residual) The point: old x always appears on the right-hand side — never wiped, only added to Addition stacks, never transforms — a new idea like “looking for a vowel” needs multiplication
microgpt_en.py · the three additionszip = two lists side by side
111	x = [t + p for t, p in zip(tok_emb, pos_emb)] # …
	# ★ x, tok_emb, pos_emb: the vector plus two notes
⋯
134	    x = [a + b for a, b in zip(x, x_residual)]
	# ★ left: the new x; right: x_residual, the old x, kept
⋯
141	    x = [a + b for a, b in zip(x, x_residual)]
	# ★ the identical line: once after attention, once after MLP

Warm-up 3a② Multiplication I: the dot product

Two vectors, 16 numbers each. This page does one thing: work out their dot product once. x = [ 0.4, 0.3, 0.5, … ] ← the vector itself, that “m at position 3” w = [ 0.2, -0.1, 0.9, … ] ← a question, for instance “is this a vowel?”   the numbers in it come from training, not from a person Dot product = multiply term by term, then add it all up 0.4×0.2 + 0.3×(−0.1) + 0.5×0.9 + … 0.08 −0.03 0.45 0.73 13 terms left 1.23 ← one number only two 16-dim vectors in, one number out In Python it is this one line sum(wi * xi for wi, xi in zip(w, x)) multiply term by term zip: two lists side by side, one (wi, xi) pair at a time sum(…) = add them all together What this number means “how well x matches this question” large positive → strong support near 0 → weak · negative → opposite One more factor: the larger the numbers in the vectors, the larger the dot product (Fig 4a) One dot product yields one number, but the main line needs 16 — so it is done 16 times.

Warm-up 3b② Multiplication II: a matrix is a stack

Python has no matrix type — a matrix is a nested list (a list of lists), one vector per row. w = [ [ 0.2, -0.1, 0.9, … ], [ 0.7, 0.3, -0.4, … ], ⋮ [-0.5, 0.8, 0.1, … ] ] ← row 0, 16 numbers ← row 1 ← row 15 Each row is itself a 16-dim vector — 16 rows in all here (how many rows is up to the matrix: fc1 has 64, lm_head 27. Note: w in Warm-up 3a was one vector; here w is a stack) One dot product of x with each row x (16 numbers) x · row 0← from Warm-up 3a 1.23 x · row 1 −1.20  ⋮ ⋮ x · row 15 0.34 as many rows, as many numbers out 16 rows here, so a 16-dim vector again In code: the outer loop runs the rows, the inner is the line above def linear(x, w): return [sum(wi * xi for wi, xi in zip(wo, x)) for wo in w] dot product with one row wo takes each row in turn(lowercase wo = one row; unrelated to the capital Wo on the right) matrix = lambda nout, nin, …: [[… for _ in range(nin)] for _ in range(nout)] Building one is a nested list too: outer nout = rows, inner nin = numbers per row Used 7 times in all, always the same Wq Wk Wv Wo ← attention fc1 fc2 ← MLP lm_head ← exit only the numbers in the rows differ

Warm-up 4a③ Non-linearity — the network's if (relu)

relu = a threshold: any signal not strong enough becomes zero −1.2 +3.4 −0.8 +0.1 relu 0 +3.4 0 +0.1 Negatives go to zero, positives pass unchanged — no if statement, just one expression Why this step has to be here Addition stacks, multiplication changes angle — both are linear Non-linearity is the only place that can tell cases apart Without relu the MLP's two linear steps collapse into one — the middle layer does nothing softmax is non-linear too, but a supporting act; “only three operations” still holds That is all three: ① addition ② multiplication (dot product / matrix) ③ non-linearity (relu) Every step that follows falls into one of these three
microgpt_en.py · non-linearityrelu's derivative: 0 or 1
50	def relu(self): return Value(max(0, self.data),
	↪     (self,), (float(self.data > 0),))
	# ★ self.data: the number itself; parens = the note
⋯
139	x = [xi.relu() for xi in x]   # all 16 pass the threshold
The note is 0 or 1 · detectors that stay off are never adjusted
	# relu's _local_grads = (1,) or (0,)
	# 0 → no gradient flows back this round (Fig 9.5)

Warm-up 4brmsnorm — the volume desk

rmsnorm = a volume desk: it sets loudness, never direction [ 3.2, -1.6, 4.8, … ] ← additions and matrices make loudness drift ÷ RMS = 4 (root mean square of the 16, illustrative) [ 0.8, -0.4, 1.2, … ] ← same direction, standard loudness It runs three times, always before a dot product before the main line L112 before attention L117 before the MLP L137 The pattern: level the loudness first, then compare against other vectors Without it, loudness keeps growing and dot-product scores lose a common scale, so similarity is misjudged and softmax lets one extreme value take nearly everything No parameters: not one of the 4,192 belongs to it — plumbing, not knowledge rmsnorm and softmax are supporting acts — “only three operations” still holds
microgpt_en.py · rmsnormthree lines, no parameters
103	def rmsnorm(x):   # ★ rmsnorm: the volume desk
104	    ms = sum(xi * xi for xi in x) / len(x)   # ms: mean loudness
105	    scale = (ms + 1e-5) ** -0.5   # scale: back to standard
106	    return [xi * scale for xi in x]   # loudness only
The three places it runs · each immediately before a linear
112	    x = rmsnorm(x)   # before the main line
⋯
117	    x = rmsnorm(x)   # before attention
⋯
137	    x = rmsnorm(x)   # before the MLP

Warm-up 5Value: the plumbing training needs

an ordinary number 0.13 a Value 0.13 .data + a note: “its sources, and the derivative for each” _children / _local_grads Forward (the answer) exactly like a plain decimal, ignore it Backward (training) following the notes backwards answers “which way each parameter must move to lower loss” (worked out by hand at training time) Value is not part of the model — it is what makes training possible. From here on it can be read as an ordinary float.
microgpt_en.py · only these linesL30–72 · a third of the file
30	class Value:
⋯
33	def __init__(self, data, children=(),
	↪              local_grads=()):
34	    self.data = data      # ★ the number itself
	# forward uses .data only
⋯
39	def __add__(self, other):
⋯
41	    return Value(self.data + other.data,
	↪                 (self, other), (1, 1))
	# ★ add: compute the sum, note it down
	#   “sources: these two; derivative 1 each”
The price of zero dependencies: the note PyTorch gives you is written out here in 40 lines.

Fig 0.5aThe first 100 lines are not the Transformer

Infrastructure — plumbing, all of it; not one line here is the Transformer L9–27 ① data + tokenizer text → integers see Fig 2 L30–72 ② Value / autograd numbers that carry notes see Warm-up 5 L75–90 ③ state_dict + params 4,192 pieces of knowledge see Fig 0 L94–106 ④ three helper functions linear/softmax/rmsnorm see Warm-up 3b · 4b L108–144 ⑤ def gpt(...) the model itself, barely thirty lines — all of this talk happens here L150–199 ⑥ training loop / inference loop loops around gpt() (see Fig 9.5, Fig 10) “A 200-line GPT” does not mean the model is 200 lines. The first 100 lines are preparation; the model itself is the thirty-odd lines in the middle — that is the part to understand.

Fig 0.5bCalling gpt(): what the four arguments are

logits = gpt( token_id, pos_id, keys, values ) token_id the number of the one letter in hand right now 0–26 pos_id the current position 0–15 keys, values the k and v left by earlier letters the only intermediate values kept across rounds (see Fig 9) returns logits = 27 scores, one per next letter (see Fig 7) ★ The one thing to make explicit In this implementation gpt() takes one token per call. The e and m before it are held in keys / values. One call = one letter in, 27 scores out; the earlier text is not in the argument list, it is in the cache.
microgpt_en.py · signaturemodel itself, L108–144
108	def gpt(token_id, pos_id, keys, values):
	# ★ gpt: the model itself, this one function
109	    tok_emb = state_dict['wte'][token_id]
110	    pos_emb = state_dict['wpe'][pos_id]
⋯
144	    return logits
	# 27 scores, nothing more
⋯
194	logits = gpt(token_id, pos_id, keys, values)
	# ★ how the outer loop calls it
The two kinds of number in Warm-up 0: keys / values are computed values, kept across steps only by the cache.

Fig 1Overview — one pass returns 27 scores

em Tokenizer char → id [e]=4 Embedding wte[id] + wpe[pos] Attention + MLP the “Block” lm_head 16 → 27 scores 27 = candidate letters Sampling weighted draw the only random choice in generation m the letter drawn becomes the next round's input (autoregression) appended each round all read back one KV cache per layer k, v of earlier letters state kept across rounds, not data flowing past Shape letter int 16 16 16 27 draw 1 letter 16 dimensions the whole way
microgpt_en.py · inference loopthe whole pipeline in 11 lines
189	for sample_idx in range(20):   # ★ sample_idx: which name
190	    keys, values = [[] for _ in range(n_layer)],
	↪         [[] for _ in range(n_layer)]
	# ★ keys, values: the KV cache itself — append only (Fig 9)
191	    token_id = BOS   # ★ token_id: the letter in hand
192	    sample = []   # ★ sample: letters produced so far
193	    for pos_id in range(block_size):
194	        logits = gpt(token_id, pos_id, keys, values)
	# ★ gpt: the whole model in one function; logits: 27 scores
195	        probs = softmax([l / temperature for l in logits])
	# ★ probs: 27 probabilities; temperature in Fig 8
196	        token_id = random.choices(range(vocab_size),
	↪             weights=[p.data for p in probs])[0]
197	        if token_id == BOS:
198	            break
199	        sample.append(uchars[token_id])

Fig 2Tokenizer — just numbering characters

Every character occurring in the training data, sorted and numbered: a 0 b 1 c 2 e 4 ⋯ z 25 + BOS 26 → vocab = 27 26 letters + 1 special BOS = Beginning of Sequence — literally “start”, but here it works as a separator emma [ 4, 12, 12, 0 ] One letter, one token — no BPE, no subwords A marker at both ends — deliberately the same symbol (one separator, used twice): BOS 26 e 4 m 12 m 12 a 0 BOS 26 ▲ “it starts” ▲ “it ends” “How does the model know when to stop?” Stopping is a token it has to learn to predict — drawing BOS ends generation. So BOS stands at both the front and the back — the very same number 26.
microgpt_en.py · tokenizerone line: sorted(set(…))
24	uchars = sorted(set(''.join(docs))) # unique characters …
	# ★ uchars: sorted character table; docs: all 32,033 names
25	BOS = len(uchars) # token id for a special Beginning …
	# ★ BOS: start-and-end symbol = 26 (hence vocab_size = 27)
26	vocab_size = len(uchars) + 1 # total number of unique …
Training · BOS wraps both ends
157	tokens = [BOS] + [uchars.index(ch) for ch in doc] + [BOS]
	# ★ tokens: one name as numbers [26, 4, 12, 12, 0, 26]
Generating · drawing BOS ends it
196	token_id = random.choices(range(vocab_size),
	↪     weights=[p.data for p in probs])[0]
197	if token_id == BOS:
198	    break

Fig 3Embedding — from symbol into meaning

token id = 4 (“e”) position pos = 1 wte (27 rows × 16 numbers) row 4 = the vector for “e” ⋮ wpe (16 rows × 16 numbers) row 1 = vector for position 1 ⋮ [ 16 numbers ] [ 16 numbers ] + addtwo meanings stacked rmsnorm ← volume desk: loudness pulled  back to standard, direction kept onto the main line (16 dims) residual stream from here on, every step only adds to this line This is where a symbol turns into meaning wpe is the only source of position — without it, emma and amme are identical. Adding them does not blur them: 16 numbers are room enough to point the two meanings in different directions, so each later matrix can still ask for the one it needs.
microgpt_en.py · first stop in gpt()two lookups + one addition
108	def gpt(token_id, pos_id, keys, values):
109	    tok_emb = state_dict['wte'][token_id] # token embedding
110	    pos_emb = state_dict['wpe'][pos_id] # position embedding
	# the two notes from earlier — here is where they come from
111	    x = [t + p for t, p in zip(tok_emb, pos_emb)] # joint …
	# addition, on stage for real
112	    x = rmsnorm(x) # note: not redundant due to backward …
	# the volume desk (rmsnorm) starts work

Fig 3bThe context limit is the row count of wpe

wpe is a matrix with only block_size = 16 rows: pos 0 [ 16 numbers ] pos 1 [ 16 numbers ] ⋮ pos 15 [ 16 numbers ] block_size = 16 (longest name = 15 letters, plus the leading BOS = 16) pos 16 ✗ no such row not found → IndexError not degraded — nonexistent Position can be supplied in two ways: Learned ordinary parameters, trained like wte GPT-2 · this project A fixed formula computed with sin/cos, never trained the original Transformer paper
microgpt_en.pywhat “128K context” actually is
77	block_size = 16 # maximum context length
	↪ of the attention window
	↪ (note: the longest name is 15 characters)
Matrix size · 16 rows, one per position
81	state_dict = {'wte': matrix(vocab_size, n_embd),
	↪ 'wpe': matrix(block_size, n_embd),
	↪ 'lm_head': matrix(vocab_size, n_embd)}
	# wpe has only block_size = 16 rows
Used as a lookup · pos_id above 15 raises IndexError
110	    pos_emb = state_dict['wpe'][pos_id] # position …
	# pos_id above 15 → this line raises IndexError

Fig 4Transformer Block — one across, one down

the x at this position (16 dims) Attention —— across “Which of the earlier information is needed?” reads across positions: the only place a token can take from others x now carries meaning from earlier tokens MLP —— down “What follows from that information?” a two-layer network inside the block; reads no other token the updated x (16 dims) These two, alternating, are the whole Transformer Moving Attention = where to take it from Refining MLP = what to make of it
microgpt_en.py · one Block in full119–132 in Figs 4a–4c
114	for li in range(n_layer):
115	    # 1) Multi-head Attention block
116	    x_residual = x
117	    x = rmsnorm(x)
118	    q = linear(x, state_dict[f'layer{li}.attn_wq'])
	# q, k, v are defined in Fig 4a-0; structure first
⋯
133	    x = linear(x_attn, state_dict[f'layer{li}.attn_wo'])
	# ★ x_attn: the four head outputs joined (16 dims, Fig 4b-0)
134	    x = [a + b for a, b in zip(x, x_residual)]
135	    # 2) MLP block
136	    x_residual = x
137	    x = rmsnorm(x)
138	    x = linear(x, state_dict[f'layer{li}.mlp_fc1'])
139	    x = [xi.relu() for xi in x]
140	    x = linear(x, state_dict[f'layer{li}.mlp_fc2'])
141	    x = [a + b for a, b in zip(x, x_residual)]

Fig 4 · timelinePositions laid out — computing the 4th

Across · reads back the first three are long since done; only the rightmost is computing BOS e m m ◀ only this  one now k,v k,v k,v q,k,v first three k,v reused as is q, k, v are defined in Fig 4a-0 — this page is about which positions update and which do not. each earlier meaning merged in, in proportion Down · this one only once merged, a single arrow remains BOS e m m ・ ・ ・ ← the first three do not move this round MLP the updated x (16) The MLP works on this one 16-dim vector — no data from other positions reaches it
microgpt_en.pyacross reads all · down works alone
Across · reads back the k, v of every earlier position
121	keys[li].append(k)
122	values[li].append(v)
⋯
129	attn_logits = [sum(q_h[j] * k_h[t][j]
	↪     for j in range(head_dim)) / head_dim**0.5
	↪     for t in range(len(k_h))]
130	attn_weights = softmax(attn_logits)
131	head_out = [sum(attn_weights[t] * v_h[t][j]
	↪     for t in range(len(v_h)))
	↪     for j in range(head_dim)]
Down · only this one vector x is available
138	x = linear(x, state_dict[f'layer{li}.mlp_fc1'])
139	x = [xi.relu() for xi in x]
140	x = linear(x, state_dict[f'layer{li}.mlp_fc2'])

Fig 4a-0Meet q, k, v — three views of one x

x after rmsnorm = what this “m” understands right now three 16×16 matrices (multiply ×3) × Wq × Wk × Wv q (16) what is being looked for a query · used once k (16) what this position is the spine label v (16) what it can supply the contents k and v go into the KV cache — one per position later cut into 4 pieces 4 each per head (Fig 4b-0) q and k only set the shares; what is taken is v. q · k set the shares v Fig 4a walks it through with real numbers: score → share → take (where x comes from: wte + wpe + earlier text — Figs 3 and 4)
microgpt_en.py · three slipsthe longest line, unfolded
Three slips · all computed from the same x
118	q = linear(x, state_dict[f'layer{li}.attn_wq'])
119	k = linear(x, state_dict[f'layer{li}.attn_wk'])
120	v = linear(x, state_dict[f'layer{li}.attn_wv'])
	# ★ q / k / v: the query / the spine label / the contents
① Score · unfolded: for each earlier t, dot q with its k, then ÷√4
129	attn_logits = [sum(q_h[j] * k_h[t][j]
	↪     for j in range(head_dim)) / head_dim**0.5
	↪     for t in range(len(k_h))]
	# ★ q_h, k_h: the 4 dims of q, k given to this head

Fig 4aOne head does only three things

Illustrative, not checkpoint weights · “e m” so far, choosing the 3rd letter · head 0 · 4 dims per vector ① Score the current q is compared once with the k of each earlier position q (query for “m” · 4 numbers) k(BOS) k(e) k(m) 2.7 0.5 −1.2 higher score = more relevant (÷√4 = ÷2: keeps scores from blowing up) ② Share softmax: squeezes the three scores into shares that add to exactly 100% 0.85 0.13 0.02 look more at one, look less at another — the total to share is fixed at 1 ③ Take 85% of v(BOS) 13% of v(e) 2% of v(m), all added up 0.85·v(BOS) + 0.13·v(e) + 0.02·v(m) head_out (4 numbers) 3 earlier letters or 300, the output is always 4 With these illustrative values, head 0 draws mainly on the BOS position.
microgpt_en.py · inside one headthree steps, one line each
⋯ 118-120: q, k, v = the three slips from Fig 4a-0
① Score
129	attn_logits = [sum(q_h[j] * k_h[t][j]
	↪     for j in range(head_dim)) / head_dim**0.5
	↪     for t in range(len(k_h))]
	# ★ attn_logits: the scores (2.7 / 0.5 / −1.2)
② Share
130	attn_weights = softmax(attn_logits)
	# ★ attn_weights: shares summing to 1 (0.85 / 0.13 / 0.02)
③ Take
131	head_out = [sum(attn_weights[t] * v_h[t][j]
	↪     for t in range(len(v_h)))
	↪     for j in range(head_dim)]
	# ★ head_out: the 4 numbers taken back in proportion

Fig 4b-0Four of them: split, run each, rejoin

Those three steps of Fig 4a actually run four times at once. q, k, v — 16 numbers each 0-3 4-7 8-11 12-15 just cut apart; nothing new head 0 ① Score ② Share ③ Take head 1 ① Score ② Share ③ Take head 2 ① Score ② Share ③ Take head 3 ① Score ② Share ③ Take each group runs those three steps once — the four never exchange anything 4 │ 4 │ 4 │ 4 concat → 16 numbers — not an element-wise sum Split, run, rejoin — those three actions and no more Not one extra parameter: Wq / Wk / Wv are still 16×16.
microgpt_en.py · multi-headsplit · run · rejoin
79	head_dim = n_embd // n_head
	# ★ head_dim: 16 ÷ 4 = 4
Split · one slice, nothing else
124	for h in range(n_head):
125	    hs = h * head_dim   # ★ hs: slice start (0, 4, 8, 12)
126	    q_h = q[hs:hs+head_dim]
127	    k_h = [ki[hs:hs+head_dim] for ki in keys[li]]
128	    v_h = [vi[hs:hs+head_dim] for vi in values[li]]
	# ki / vi: the k, v of each earlier position in the cache
Run · 129-131 = exactly the three steps of Fig 4a
Rejoin · [4 │ 4 │ 4 │ 4] concat → 16, not a sum
132	    x_attn.extend(head_out)
	# ★ x_attn: the 16 numbers from the four heads joined

Fig 4bWhy four heads: four patterns at once

Illustrative, not checkpoint weights · same moment, same layer: each head forms its own pattern BOS e m m head 0 on the previous letter head 1 on the start — how far into the name head 2 on the first letter head 3 spread out — overall length (one head could only pick one of these) One head can look at many positions, but only with one pattern A dot product already sums four groups and then adds those sums together. Splitting into heads = skipping that last addition, so each group shares on its own. One head: add it all first, then share one set of shares → every feature uses it Four heads: no final add; each shares four sets of shares → four subspaces decide
microgpt_en.py · multi-headthis figure = that for loop ×4
No new code · split, run, rejoin from Fig 4b-0, four full rounds
124	for h in range(n_head):
⋯
132	    x_attn.extend(head_out)

Fig 4cAttention wiring — how four heads return

main line x (16) skip: passes by untouched rmsnorm volume desk: loudness back to standard, direction kept Wq Wk Wv q(16) k(16) v(16) k and v written in KV cache history of k, v k, v of earlier positions kept, not recomputed Split, run, rejoin = Fig 4b-0 each of the four heads runs those three steps, sharing nothing [4 │ 4 │ 4 │ 4] concat → 16 read KV Wo ← the only place the four heads mix + residual: what was computed is added back to x — that is what the bypass is for; x is never overwritten on to the MLP
microgpt_en.py · the attention block124–132 in Fig 4b-0
116	x_residual = x
117	x = rmsnorm(x)
118	q = linear(x, state_dict[f'layer{li}.attn_wq'])
119	k = linear(x, state_dict[f'layer{li}.attn_wk'])
120	v = linear(x, state_dict[f'layer{li}.attn_wv'])
121	keys[li].append(k)   # this position's k, v stored
122	values[li].append(v)
123	x_attn = []
124	for h in range(n_head):
125	    hs = h * head_dim
126	    q_h = q[hs:hs+head_dim]   # the “split”, one line
⋯
132	    x_attn.extend(head_out)
133	x = linear(x_attn, state_dict[f'layer{li}.attn_wo'])
	# Wo: the only place the four heads mix
134	x = [a + b for a, b in zip(x, x_residual)]
	# the bypass rejoins here

Fig 5MLP — 64 soft detectors and a threshold

main line x (16) passes by untouched rmsnorm fc1 (16 → 64) 64 “detectors” each row dots with x: a soft “is X present?” −1.2 +3.4 −0.8 +0.1 ⋯ relu ── the threshold negative → 0 (off) positive → strength kept 0  +3.4  0  +0.1 ⋯ fc2 (64 → 16) 64 “responses” fc2 combines them into a 16-dim correction + One detector might ask “two consonants in a row?” once it fires, it adds “a vowel is due” to the line relu stops the two linear steps merging, so behaviour depends on the input
microgpt_en.py · MLPwider → filter → narrower
All of relu · the derivative is 0 or 1
50	def relu(self): return Value(max(0, self.data),
	↪     (self,), (float(self.data > 0),))
The two 4× matrices · 64 soft feature detectors
87	state_dict[f'layer{i}.mlp_fc1'] = matrix(4 * n_embd, n_embd)
88	state_dict[f'layer{i}.mlp_fc2'] = matrix(n_embd, 4 * n_embd)
	# fc1 = 64 detector scores; fc2 writes back to 16
The MLP block itself · three lines + the residual
136	x_residual = x
137	x = rmsnorm(x)
138	x = linear(x, state_dict[f'layer{li}.mlp_fc1'])
139	x = [xi.relu() for xi in x]
140	x = linear(x, state_dict[f'layer{li}.mlp_fc2'])
141	x = [a + b for a, b in zip(x, x_residual)]

Fig 6Residual — every sub-layer adds back

No stage replaces the main line — each correction is computed from the latest x, then added back Any stage may abstain (output ≈ 0 → everything passes through), so more layers never hurt embedding lm_head a 16-lane bus read + added back attn0 computes a correction read + added back mlp0 computes a correction Every stage: read the bus → compute a correction → add it back, never overwrite Written out, the vector entering lm_head is literally: x_final = the normalised embedding (letter + position)     + the attention correction     + the MLP correction More layers just insert two more plus signs on this bus — the architecture is untouched
microgpt_en.pyevery plus in the talk is this line
The (1, 1) here · the note addition writes — used in Fig 9.5
39	def __add__(self, other):   # ★ other: the other addend
40	    other = other if isinstance(other, Value)
	↪         else Value(other)
41	    return Value(self.data + other.data,
	↪         (self, other), (1, 1))
	# the note says (1, 1) — gradients pass through unchanged
A whole Block = two plus signs on a loop
114	for li in range(n_layer):   # more layers = one more round
⋯
134	    x = [a + b for a, b in zip(x, x_residual)]
⋯
141	    x = [a + b for a, b in zip(x, x_residual)]

Fig 7lm_head — a vector back into a letter

final x on the main line (16) a dot product with each of the 27 candidate rows (⊙) each row is 16-dim; each yields one logit row for “a” ⊙ +1.8 row for “e” ⊙ +0.4 row for “m” ⊙ +3.1 ⋮ row for “z” ⊙ −2.0 whichever matches the current x best gets the highest logit — reading out, not thinking Back to the dot product of Warm-up 3a: its size depends on two things ① the directions  ② the scale on both sides Scale amplifies negative scores as well as positive; it is no fixed bias or head start It no longer reasons — it only outputs Purely linear · one position · no relu · no other tokens. Translating needs only a dictionary. Because it is this simple, everything before must do the work — all pressure comes through here.
microgpt_en.py · the exitstructurally unable to reason
Same shape (27 × 16) · opposite directions, independent weights
81	state_dict = {'wte': matrix(vocab_size, n_embd),
	↪ 'wpe': matrix(block_size, n_embd),
	↪ 'lm_head': matrix(vocab_size, n_embd)}
	# lm_head has wte's shape but not its weights
All of the “translation” · one line of dot products, no relu
94	def linear(x, w):
95	    return [sum(wi * xi for wi, xi in zip(wo, x)) for wo in w]
⋯
143	logits = linear(x, state_dict['lm_head'])
	# the last matrix: 27 candidate rows, one score each
144	return logits

Fig 8Sampling — the only random choice here

27 logits ÷ temperature 0.5 → wider gaps · safer 1.5 → flatter · more random÷0.5 doubles the gaps (1.3 apart → 2.6)softmax reads the gap, not the ratio softmax squeezes 27 scores into shares adding to exactly 100% one weighted draw the only randomness here one letter With only BOS — what the model expects first (illustrative) a m j k ⋮ x ← almost no name starts with x After seeing e, m, m, a BOS “the name is over” — very high probability n l Vowels versus consonants, name length — nobody specified any of it This distribution is a by-product of guessing the next letter — that is what learning looks like
microgpt_en.py · samplingrandomness only in choices
Temperature · just dividing the scores
187	temperature = 0.5 # in (0, 1], control the …
The draw · weighted by the shares softmax produced
193	for pos_id in range(block_size):
194	    logits = gpt(token_id, pos_id, keys, values)
195	    probs = softmax([l / temperature for l in logits])
	# temperature just divides the scores
196	    token_id = random.choices(range(vocab_size),
	↪         weights=[p.data for p in probs])[0]
	# the only random choice in generation, this one line
197	    if token_id == BOS:
198	        break

Fig 9Autoregression — putting it all together

round 1 round 2 round 3 round 4 round 5 BOS e m m a pos 0 pos 1 pos 2 pos 3 pos 4 GPT GPT GPT GPT GPT e m m a BOS BOS drawn → stop the letter drawn becomes the next round's input The KV cache grows each round — what is already there is never recomputed: round 1 k₀ v₀ round 2 k₀ v₀ k₁ v₁ round 3 k₀ v₀ k₁ v₁ k₂ v₂ round 4 k₀ v₀ k₁ v₁ k₂ v₂ k₃ v₃ round 5 k₀ v₀ k₁ v₁ k₂ v₂ k₃ v₃ k₄ v₄ new (rightmost only) reused as is Final output: emma each forward returns scores; sampling picks one token, which goes into the next round old k/v are not recomputed, yet the new q still scores every position in the cache
microgpt_en.py · one namethe Fig 1 flow, state in focus
Only the new k, v are appended each round (inside gpt())
121	keys[li].append(k)   # only the new one, each round
122	values[li].append(v)
One letter per round · the cache resets for every name
190	keys, values = [[] for _ in range(n_layer)],
	↪     [[] for _ in range(n_layer)]
	# reset for every name
191	token_id = BOS
192	sample = []
193	for pos_id in range(block_size):
194	    logits = gpt(token_id, pos_id, keys, values)
195	    probs = softmax([l / temperature for l in logits])
196	    token_id = random.choices(range(vocab_size),
	↪         weights=[p.data for p in probs])[0]
197	    if token_id == BOS:
198	        break
199	    sample.append(uchars[token_id])

Fig 9.5From generating back to learning

① Forward shown on a toy expression of just two steps w = 2 × x = 3 y = 6 ← intermediate y = 6 + c = 1 L = 7 ← treat as the loss As it computes, each step leaves a note (_local_grads): multiply → (3, 2) = the other operand: move w by 1 and y moves by 3 add → (1, 1) = move either by 1 and L moves by 1 ② Backward start at L and multiply the notes back L.grad = 1 start: L against itself, ×1 c.grad = 1 × 1 = 1 through the add note (1, 1) y.grad = 1 × 1 = 1 through the add note (1, 1) w.grad = 1 × 3 = 3 through the multiply note (3, 2) x.grad = 1 × 2 = 2 through the multiply note (3, 2) Each node does one thing: the note × its own grad, handed to its sources. Done. w.grad = 3 means: move w from 2 to 2.001 and L goes from 7 to 7.003. That number is the gradient. To lower L, nudge the other way — training is no more than this. In the real model, one backward walk over the same graph accumulates gradients for all 4,192 parameters. Add notes are all 1, and ×1 passes through unchanged — that is the residual highway of Fig 6.
microgpt_en.py · the notesheld back until now
The note each operation writes · the value in the parentheses
39	def __add__(self, other):   # a + b
41	    return Value(self.data + other.data,
	↪     (self, other), (1, 1))
	# ★ add: 1 on both sides
43	def __mul__(self, other):   # a * b
45	    return Value(self.data * other.data,
	↪     (self, other), (other.data, self.data))
	# ★ multiply: the note holds the other operand
50	def relu(self): … (float(self.data > 0),)
	# ★ relu: 1 or 0 — a closed path passes nothing
The core of backward() is three lines · L59–72
69	self.grad = 1   # ★ start from the loss
70	for v in reversed(topo):   # one pass, back to front
71	    for child, local_grad in
	↪         zip(v._children, v._local_grads):
72	        child.grad += local_grad * v.grad
	# ★ the table above = this line, run five times
topo = the nodes ordered first, so that on the way back every upstream grad is already final.

Fig 10Training — one goal: guess the next letter

One training example: “emma” · every question is fed the real previous letter, never the model's own draw BOS e m m a BOS ? ? ? ? ? One computation per position: “what is the next letter?” one “emma” = 5 learning signals, not 1 loss = −log P(correct answer) computed at each position, then averaged lower odds, heavier penalty 1 → 0 0.01 → 4.6 loss.backward() one backward pass over the same graph → 4,192 gradients Adam — a small step for each of the 4,192 avg of the gradient ÷ √avg of its square then on to the next training step Training has exactly one task — guess the next letter The chain rule runs through lm_head, MLP, attention and embedding, adjusting every parameter. Vowels and consonants, name length — all of it is a by-product.
microgpt_en.py · one training stepnothing exotic about Adam
156	doc = docs[step % len(docs)]   # ★ step: which step
157	tokens = [BOS] + [uchars.index(ch) for ch in doc] + [BOS]
158	n = min(block_size, len(tokens) - 1)   # ★ n: how many signals
⋯
163	for pos_id in range(n):
164	    token_id, target_id = tokens[pos_id],
	↪         tokens[pos_id + 1]
	# ★ target_id: the correct answer = the next letter
165	    logits = gpt(token_id, pos_id, keys, values)
166	    probs = softmax(logits)
167	    loss_t = -probs[target_id].log()
	# ★ loss_t / loss: the error (−log of the correct prob)
168	    losses.append(loss_t)
169	loss = (1 / n) * sum(losses) # final average loss …
⋯
172	loss.backward()   # which way each parameter lowers the loss
⋯
175	lr_t = learning_rate * (1 - step / num_steps) # …
	# ★ lr_t / num_steps: step size decays / 1,000 steps total
176	for i, p in enumerate(params):   # ★ p: all 4,192, one by one
177	    m[i] = beta1 * m[i] + (1 - beta1) * p.grad
	# ★ m / v: Adam's two statistics (unrelated to attention v)
178	    v[i] = beta2 * v[i] + (1 - beta2) * p.grad ** 2
⋯
181	    p.data -= lr_t * m_hat / (v_hat ** 0.5 + eps_adam)
182	    p.grad = 0
The whole Transformer in three lines
Attention moves decides where to take information from
MLP refines decides what to make of what it took
Residual adds both only add to the main line, never overwrite
BOS Tokenizer Embedding Attention + MLP lm_head Samplingrandomness a autoregression: the letter drawn becomes the next round's input the three lines above
Training: the correct next letter → loss → backward → Adam → update all 4,192 parameters
More layers = the same thing again · Today's large models = the same thing, orders of magnitude bigger
4,192 numbers · 200 lines of pure Python — addition, multiplication and one threshold, start to finish
Extras · further topics

Three directions

What the real weights look like

Inspecting a trained checkpoint: the lengths and angles of wpe, and what all four heads actually attend to at one moment.

scripts/dump_position.py
scripts/dump_attention.py

Switching to a Chinese dataset

One character per token pushes the vocabulary past 700 while block_size shrinks — the same code, a completely different split of parameters.

microgpt.py (names from Jin Yong novels)

What longer training does

The loss keeps falling, but the model starts memorising: generated names land exactly on the training data — overfitting made concrete.

EXPERIMENTS.md

Q & A