Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Learn · 1 of 7

What a language model actually is

A next-token predictor, and why that description is both accurate and far too small.

A large language model (A model trained to predict the next token over enormous amounts of text, which turns out to be enough to do a great deal.) is, mechanically, a thing that predicts the next token. Shown a sequence of text, it produces a probability distribution over what comes next, and one is drawn from it. Then it does that again, with the new token included, until it stops.

That is the whole mechanism. It is accurate, and it badly under-describes what the result can do.

Why that is strange

Nobody designed in the ability to hold a conversation, write working code, or follow an instruction it had never seen. Those appeared, as the models and the training data got bigger, out of a system that was only ever being asked to guess the next fragment of text.

This is genuinely one of the more surprising empirical results of the last decade, and it is worth not getting over too quickly. There is no module inside that handles reasoning. There is no database of facts. There is a very large function that maps a sequence to a probability distribution, trained until the predictions got good.

Why that explains the failures too

Keeping the mechanism in mind is the cheapest way to predict how these things go wrong.

A model does not know whether it knows something, because nothing in the mechanism distinguishes recall from generation. Asked for a citation it has not seen, it produces the most plausible-looking one — correct-shaped authors, sensible title, a year that fits. This is not lying, and it is not a bug awaiting a patch; it follows directly from what the thing is.

It also explains why showing beats telling. The model is predicting a continuation, so two worked examples of the output you want shape the prediction far more than a paragraph of adjectives describing it.

What comes next

Everything else in this section builds on this one. Tokens is the unit it actually works in, and almost every cost and limit on this site is counted in them.

How a model gets made

Width shows roughly where the compute goes

The four stages of training a modelFour stages left to right — pretraining, fine-tuning, preference tuning, evaluation and release — drawn as bands whose widths show roughly where the compute goes. Pretraining dwarfs the rest. Each stage is described in the list below.1months, most of the moneyPretraining2daysFine-tuning3daysPreference tuning4ongoingEvaluation and releaseYou are never charged for any of this. What you pay for is inference — running the finished model.
  1. Pretraining (months, most of the money). Predict the next token across an enormous body of text. This is where almost everything the model knows comes from, and almost all of the cost.
  2. Fine-tuning (days). Train on curated example conversations. This is what turns a text completer into something that answers a question.
  3. Preference tuning (days). Train against human judgements of which answer is better. Tone, helpfulness and refusals come from here.
  4. Evaluation and release (ongoing). Score it, decide whether to ship it, and keep scoring it afterwards.
ShareOpen LinkedIn