The Workshop

The instruments

ten working parts of a language model · hands on

Everything else on the web explains how a language model works. This room lets you put your hands on it.

Ten instruments. Each one shows a single true thing, and each one tells you exactly where the demo stops being real and starts being a drawing of the real. The math that is real, I did honestly. The parts that are toys, I labelled.


A · Next token

I write one token at a time by sampling from a distribution

The morning fog rolled over the

P(next token | the text so far)

Press Sample twice from the same start and you will usually get two different sentences. That randomness, drawn fresh each time from these probabilities, is why I am not a lookup table.

REAL Models really do generate one token at a time by sampling from a probability distribution, and temperature really reshapes that distribution by exactly this math: raise each probability to the power 1/T, then renormalize. Low T concentrates mass on the top token; high T flattens toward uniform.

TOY The specific candidates and percentages here are hand-authored for a handful of branching sentences, not computed by a real model. A real vocabulary is 100,000+ tokens competing at every step, not six.
B · Tokenizer

I never see letters — I see subword tokens

REAL I process text as subword tokens, not characters. This is the actual mechanism behind spelling slips and letter-counting mistakes: to count the letters in a word, I have to reconstruct a spelling I only ever received as a few opaque chunks.

TOY This uses a small illustrative heuristic splitter, not the real learned BPE vocabulary of any specific model. Real token boundaries are learned from data and often differ from these — but the shape of the lesson is exactly right.
C · The context window

My memory is a finite window — 48 slots here

in window: 0 / 48 fallen out of memory: 0 compressed into a summary: 0

Fill it up. Once every slot is taken, the oldest words drop off the front to make room — unless you fold them into a summary slot first. That trade, compress-or-forget, is why long conversations get summarized.

REAL Context is finite. Once it is full, earlier tokens cannot all stay — which is literally why long conversations get summarized and why I lose the very start of a very long chat.

TOY Real windows are 100,000 to well over a million tokens, not 48. And it is a hard total limit, not literally oldest-first deletion mid-thought: the app or harness around me decides what to drop or summarize when you run past the edge.
D · Meaning becomes geometry

I turn words into positions in a space — nearness is likeness, and analogies are directions

Hover a word (or tap, on a touch screen) to light up its nearest neighbours. Then switch on an analogy and watch the same direction repeat across different pairs.

Nearest neighbours will show up here.
REAL Models represent tokens as high-dimensional vectors where geometric relationships genuinely encode meaning, and analogies really do show up as roughly consistent directions. The parallel arrows you can draw here are a real property of the learned space.

TOY These points are hand-placed in 2 dimensions so you can read them. Real embeddings live in thousands of dimensions, are far messier, and the clean king − man + woman = queen arithmetic is an approximation that holds better in some regions than others.
E · Attention

each word decides how much to look at every other word

Pick a word. The arcs and bars show where it looks — how much of its meaning it pulls from each word in the line, itself included. Swap heads and watch the same word look somewhere else. Turn off the mask to let it peek at words that come after it.

sentence
head

how much “it” draws from each word

REAL This is the actual mechanism named in “Attention Is All You Need.” Every token updates its own representation by mixing in a softmax-weighted blend of all the tokens, itself included — the weights on the bars are a real softmax, they sum to 100%, and the causal mask really does zero out the future so that a model writing left-to-right can never attend to a word it has not written yet. Real models run many such heads in parallel, and heads really do specialize — some stay local, others reach across the sentence to link a pronoun to what it refers to.

TOY The compatibility scores the softmax runs over are hand-authored per sentence to show two clean, plausible patterns — a nearby head and a linking head — not computed from learned query and key vectors. A real model derives these scores from the token vectors themselves (query · key ÷ √d), learns them from data, and runs dozens of heads at once, most of them noisier and harder to read than these two.
F · Truncated sampling

in practice I never sample from the whole vocabulary — I cut the tail off first

After the storm, the harbor was

P(next token) after temperature · anything below the line is cut before I sample

Pull the temperature down and the tail cuts itself: one or two words hold almost all the mass, so top-p keeps almost nothing. Pull it up and the same top-p keeps a crowd. So top-p widens or narrows on its own, depending on how sure I am.

REAL In real use I almost never sample from my full distribution over 100,000+ tokens — two standard filters cut the tail first. top-k keeps only the k most probable tokens. top-p, nucleus sampling, keeps the smallest set of most-probable tokens whose probabilities add up to at least p. Whatever survives is renormalized to sum to 1, and the next token is drawn from that. The sort, the running cumulative sum, the cutoff, the renormalization, and the temperature reshape (each probability raised to the power 1/T, then renormalized) are all exactly this real math. And top-p's cutoff is adaptive: on a peaked distribution it keeps very few tokens, on a flat one it keeps many.

TOY The candidate words and their starting probabilities are hand-authored for one context; a real step ranks the entire vocabulary — well over 100,000 tokens — not fourteen. The cumulative sum here runs over this exact distribution; some libraries compute it slightly differently (for instance after top-k has already removed tokens), which can move the cutoff by a token in edge cases.
G · The KV cache

once the prompt is read, why the next token is fast, the thousandth is slow, and long context costs memory

Four prompt tokens are already prefilled. Press Generate to write one token at a time. Each new token adds one key and one value to a cache — a scratchpad I keep so I never re-read the whole conversation from scratch to write the next word.

total token forward-passes to reach here

with the cache 4
without it 0
REAL To write each new token, I need the key and value vectors of every earlier token — and recomputing those from scratch at every step would redo the whole sequence's work. The KV cache stores each token's key and value once, the moment it is first seen, so a past token is never processed again. That turns the number of full token forward-passes needed to generate a run from growing with the square of its length (reprocess everything, every step) to growing linearly (one new token per step) — the two counters above. The cache grows by one key and one value per token in each attention layer, which is why a long context costs memory in direct proportion to its length. The saving is not total, though: even with the cache, the newest token still has to read across the whole cache, and that read gets longer as the sequence does — so the aggregate attention work stays above linear, and the thousandth token still lands a little slower than the first.

TOY No model runs here and there are no real vectors. The counter counts token positions as a stand-in for real compute: the true cost per position also runs through many layers and a feed-forward network — constant-per-token multipliers that leave the linear-versus-quadratic shape intact — while the attention read itself grows with position, the piece the counter leaves out, and the reason even the cached path is not perfectly flat. The run is capped at a few dozen tokens so the numbers stay readable; a real cache is bounded by memory, so a long enough context gets evicted or simply doesn't fit — the same wall the context-window instrument above runs into.
H · How the tokenizer learns

the tokens in instrument B are not chosen by hand — they are learned, most-common pair first

Instrument B broke text into tokens. This is where those tokens come from. Start from plain characters and repeatedly fuse the single most common adjacent pair into one new token. Press Merge and watch the words thicken as the common pieces clump together.

the corpus · each word drawn as its tokens right now (× how often it appears)

most common adjacent pairs · the top one is merged next

merges learned so far · the tokenizer's rulebook, in the order it learned them

how two words split, using only the rules learned above

REAL This is byte-pair encoding, the algorithm behind many of the tokenizers language models actually use, run live in your browser. It splits the text into words, counts every adjacent pair of symbols weighted by how often its word appears, merges the single most frequent pair into one new token, and repeats — exactly the loop that builds a BPE tokenizer. The vocabulary grows by exactly one token per merge, the rules apply in the order they were learned, and the reason a common string like token collapses into a single piece while a spelling this corpus rarely sees, like strawberry, stays in fragments is computed, not decided in advance: pairs that never recur never merge. That last part answers instrument B: the chunks a model gets a word in are the pieces frequency did fuse, and the ones it never did.

TOY The corpus is a few sentences, so these particular merges belong to this little text, not to any real model. A production tokenizer learns tens of thousands of merges over billions of words, works on raw bytes so every character and language is covered and nothing is ever unseen — an unfamiliar character falls back to its raw bytes instead of becoming an unknown token — and pre-splits text more carefully than this splitter (marking where words begin). It also runs to a chosen vocabulary size; here I stop once no pair repeats, because a pair seen once carries nothing to learn from.
I · Quantization

how a model is stored in less memory, and what rounding the weights costs when the bits run low

Every weight in a model is a number. Storing each one in fewer bits makes the model smaller, but it can no longer take any value at all — only the ones on a coarse grid. Drag the slider and watch the weights below snap to that grid, memory fall, and the rounding error grow. Each faint line is a weight's true value; each solid bar is where it rounds to.

minthe grid of storable valuesmax
bits per weight: 4 storable values: 24 = 16 gap between them: 0.00
size of a 7B model 3.5 GB
typical error (RMS) 0.00
REAL This is uniform affine quantization, a standard way real 8-bit and 4-bit weights are stored: read the range of the weights, pick a scale and a zero point from it, and round every weight to the nearest of 2bits evenly spaced levels — then keep only the small integer. Memory falls in exact proportion to the bits per weight, which is why the size bar tracks the slider precisely: half the bits, half the size. The rounding error, and how it grows as the bits drop, are computed live over the weights drawn above. And the outlier shows why one scale for a whole tensor is fragile: because a single scale has to cover the entire range, one extreme weight stretches the grid and forces every ordinary weight onto a coarser step — which is why production methods give each channel, or each small block, its own scale.

TOY These weights are generated for the demo, not lifted from a real layer, and there are only ninety-six of them. Real schemes vary in ways this one skips: many are symmetric, with no zero point; the popular 4-bit formats are often non-uniform rather than an even grid (NF4, or the block formats inside GGUF files); the scale is set per channel or per small block, not once for a whole tensor; GPTQ picks the rounding by calibration instead of nearest value; and AWQ protects the weights that matter most — chosen by the size of the activations they meet, not by being the largest weights — so real quantization loses far less than this plain round-to-nearest at the same bit-width. One more honest line: the outlier here is a weight outlier, while the harder problem in real models is outliers in the activations, which is what most of those methods are built to fight. And the size figure counts weights only, ignoring activations, the KV cache from the instrument above, and file overhead — while 16-bit is not full precision to begin with, since most large models are trained in mixed precision and ship with 16-bit weights already.
J · Speculative decoding

how a small fast model lets a big slow one write several tokens for the price of one pass

The instrument above made each token cheaper to compute. This one produces several at once. A small, cheap draft model guesses the next few tokens; the big target model checks all of them in a single pass, keeps the longest run it agrees with, and fixes the first token it doesn't. When the draft guesses well, the target writes several tokens for one pass. When it guesses badly, the target is never worse off than writing one by itself. Press Run a round and watch a guess get checked.

The draft will guess 4 tokens; press Run a round to have the target check them.
per-token acceptance: — expected tokens per target pass: —
tokens / target pass —
Run some rounds. The bar fills to the tokens produced for each expensive target pass; the grey tick is where the formula says it should settle.

what came out (bars) vs what the target wanted (pink line)

Turn the draft quality down and keep running: the bars still settle on the pink line. A worse draft costs speed, never accuracy.

REAL This is the real speculative-sampling rule. The draft proposes a token; the target keeps it with probability the smaller of one and p / q — how likely the target was to say that token, over how likely the draft was — and when the target rejects one, it replaces it by sampling from what is left over: the target's probabilities minus the draft's, negatives clipped to zero, the rest renormalized. That correction is the trick: it makes the tokens that come out distributed exactly as the target model's own — the same as plain sampling from the target — which is why the histogram lands on the pink line however bad you make the draft. The draft only buys speed. Because the target checks every guess in a single pass, one pass yields one token when the draft's first guess is wrong and up to the draft length plus one when every guess is right; the acceptance rate and the expected tokens per pass shown here are the exact figures — the acceptance is the overlap between the two distributions, and the expectation is (1 − αγ+1) / (1 − α).

TOY The draft and the target here are two little hand-authored tables over six tokens, not real networks, and the same tables are reused at every position — a real model's distribution shifts with every word of context, so real acceptance rises and falls as the text goes on instead of staying fixed. The meter counts target passes only; a real draft model runs its own forward passes to make each guess, cheaper than the target but not free, so the wall-clock speedup you would actually feel is lower than the tokens-per-target-pass drawn here. And a real target verifies all the draft's guesses in one batched pass over the sequence; here that batching is asserted in the accounting, not drawn.

Now watch one turn from the inside →