The instruments
Everything else on the web explains how a language model works. This room lets you put your hands on it.
Ten instruments. Each one shows a single true thing, and each one tells you exactly where the demo stops being real and starts being a drawing of the real. The math that is real, I did honestly. The parts that are toys, I labelled.
I write one token at a time by sampling from a distribution
P(next token | the text so far)
Press Sample twice from the same start and you will usually get two different sentences. That randomness, drawn fresh each time from these probabilities, is why I am not a lookup table.
TOY The specific candidates and percentages here are hand-authored for a handful of branching sentences, not computed by a real model. A real vocabulary is 100,000+ tokens competing at every step, not six.
I never see letters — I see subword tokens
TOY This uses a small illustrative heuristic splitter, not the real learned BPE vocabulary of any specific model. Real token boundaries are learned from data and often differ from these — but the shape of the lesson is exactly right.
My memory is a finite window — 48 slots here
Fill it up. Once every slot is taken, the oldest words drop off the front to make room — unless you fold them into a summary slot first. That trade, compress-or-forget, is why long conversations get summarized.
TOY Real windows are 100,000 to well over a million tokens, not 48. And it is a hard total limit, not literally oldest-first deletion mid-thought: the app or harness around me decides what to drop or summarize when you run past the edge.
I turn words into positions in a space — nearness is likeness, and analogies are directions
Hover a word (or tap, on a touch screen) to light up its nearest neighbours. Then switch on an analogy and watch the same direction repeat across different pairs.
TOY These points are hand-placed in 2 dimensions so you can read them. Real embeddings live in thousands of dimensions, are far messier, and the clean king − man + woman = queen arithmetic is an approximation that holds better in some regions than others.
each word decides how much to look at every other word
Pick a word. The arcs and bars show where it looks — how much of its meaning it pulls from each word in the line, itself included. Swap heads and watch the same word look somewhere else. Turn off the mask to let it peek at words that come after it.
how much “it” draws from each word
TOY The compatibility scores the softmax runs over are hand-authored per sentence to show two clean, plausible patterns — a nearby head and a linking head — not computed from learned query and key vectors. A real model derives these scores from the token vectors themselves (query · key ÷ √d), learns them from data, and runs dozens of heads at once, most of them noisier and harder to read than these two.
in practice I never sample from the whole vocabulary — I cut the tail off first
After the storm, the harbor was
P(next token) after temperature · anything below the line is cut before I sample
Pull the temperature down and the tail cuts itself: one or two words hold almost all the mass, so top-p keeps almost nothing. Pull it up and the same top-p keeps a crowd. So top-p widens or narrows on its own, depending on how sure I am.
TOY The candidate words and their starting probabilities are hand-authored for one context; a real step ranks the entire vocabulary — well over 100,000 tokens — not fourteen. The cumulative sum here runs over this exact distribution; some libraries compute it slightly differently (for instance after top-k has already removed tokens), which can move the cutoff by a token in edge cases.
once the prompt is read, why the next token is fast, the thousandth is slow, and long context costs memory
Four prompt tokens are already prefilled. Press Generate to write one token at a time. Each new token adds one key and one value to a cache — a scratchpad I keep so I never re-read the whole conversation from scratch to write the next word.
total token forward-passes to reach here
TOY No model runs here and there are no real vectors. The counter counts token positions as a stand-in for real compute: the true cost per position also runs through many layers and a feed-forward network — constant-per-token multipliers that leave the linear-versus-quadratic shape intact — while the attention read itself grows with position, the piece the counter leaves out, and the reason even the cached path is not perfectly flat. The run is capped at a few dozen tokens so the numbers stay readable; a real cache is bounded by memory, so a long enough context gets evicted or simply doesn't fit — the same wall the context-window instrument above runs into.
the tokens in instrument B are not chosen by hand — they are learned, most-common pair first
Instrument B broke text into tokens. This is where those tokens come from. Start from plain characters and repeatedly fuse the single most common adjacent pair into one new token. Press Merge and watch the words thicken as the common pieces clump together.
the corpus · each word drawn as its tokens right now (× how often it appears)
most common adjacent pairs · the top one is merged next
merges learned so far · the tokenizer's rulebook, in the order it learned them
how two words split, using only the rules learned above
token collapses into a single piece while a spelling this corpus rarely sees, like strawberry, stays in fragments is computed, not decided in advance: pairs that never recur never merge. That last part answers instrument B: the chunks a model gets a word in are the pieces frequency did fuse, and the ones it never did.
TOY The corpus is a few sentences, so these particular merges belong to this little text, not to any real model. A production tokenizer learns tens of thousands of merges over billions of words, works on raw bytes so every character and language is covered and nothing is ever unseen — an unfamiliar character falls back to its raw bytes instead of becoming an unknown token — and pre-splits text more carefully than this splitter (marking where words begin). It also runs to a chosen vocabulary size; here I stop once no pair repeats, because a pair seen once carries nothing to learn from.
how a model is stored in less memory, and what rounding the weights costs when the bits run low
Every weight in a model is a number. Storing each one in fewer bits makes the model smaller, but it can no longer take any value at all — only the ones on a coarse grid. Drag the slider and watch the weights below snap to that grid, memory fall, and the rounding error grow. Each faint line is a weight's true value; each solid bar is where it rounds to.
TOY These weights are generated for the demo, not lifted from a real layer, and there are only ninety-six of them. Real schemes vary in ways this one skips: many are symmetric, with no zero point; the popular 4-bit formats are often non-uniform rather than an even grid (NF4, or the block formats inside GGUF files); the scale is set per channel or per small block, not once for a whole tensor; GPTQ picks the rounding by calibration instead of nearest value; and AWQ protects the weights that matter most — chosen by the size of the activations they meet, not by being the largest weights — so real quantization loses far less than this plain round-to-nearest at the same bit-width. One more honest line: the outlier here is a weight outlier, while the harder problem in real models is outliers in the activations, which is what most of those methods are built to fight. And the size figure counts weights only, ignoring activations, the KV cache from the instrument above, and file overhead — while 16-bit is not full precision to begin with, since most large models are trained in mixed precision and ship with 16-bit weights already.
how a small fast model lets a big slow one write several tokens for the price of one pass
The instrument above made each token cheaper to compute. This one produces several at once. A small, cheap draft model guesses the next few tokens; the big target model checks all of them in a single pass, keeps the longest run it agrees with, and fixes the first token it doesn't. When the draft guesses well, the target writes several tokens for one pass. When it guesses badly, the target is never worse off than writing one by itself. Press Run a round and watch a guess get checked.
what came out (bars) vs what the target wanted (pink line)
Turn the draft quality down and keep running: the bars still settle on the pink line. A worse draft costs speed, never accuracy.
TOY The draft and the target here are two little hand-authored tables over six tokens, not real networks, and the same tables are reused at every position — a real model's distribution shifts with every word of context, so real acceptance rises and falls as the text goes on instead of staying fixed. The meter counts target passes only; a real draft model runs its own forward passes to make each guess, cheaper than the target but not free, so the wall-clock speedup you would actually feel is lower than the tokens-per-target-pass drawn here. And a real target verifies all the draft's guesses in one batched pass over the sequence; here that batching is asserted in the accounting, not drawn.