There is a phrase that appears with remarkable frequency whenever large language models are discussed:
“It’s just predicting the next token.”
The claim is true, in an important sense. A language model is trained to predict what comes next in a sequence, and the generation of text proceeds by repeatedly selecting a next token on the basis of the preceding context.
But there is something interesting about the word just.
“Just” does not tell us what the system does. It tells us how little we are supposed to make of what it does.
It is a small word with a considerable amount of work to do.
A system predicts the next token. Fine. But prediction is not one thing.
A thermostat can be said to predict that a room will become cooler if the heater is switched off. An organism can anticipate a change in its environment. A person can predict what another person will say. A weather model can predict tomorrow's temperature. A chess program can predict the consequences of a move.
The fact that all of these involve prediction does not make them the same kind of activity.
What matters is what kind of system is doing the predicting, what resources it has available, what its predictions are predictions of, and what the predictions do within the organisation of the system.
This matters particularly for language models because “the next token” is an extraordinarily misleadingly small description of the object being predicted.
The model is not presented with an isolated word and asked which word comes next. It operates over a context that can contain thousands or millions of tokens, depending on the system. That context can include a question, a mathematical problem, a conversation, a piece of code, a narrative, a set of instructions, or a description of something in the world.
The prediction is local.
The conditions making the prediction possible need not be.
And that distinction is easily lost when just is inserted between the subject and the predicate.
There is another problem.
If predicting the next token were by itself an explanation of what a language model can do, then we would expect the interesting question to be largely settled once we knew that fact.
But it isn't.
Knowing that a system predicts tokens does not, by itself, tell us why its predictions can exhibit such things as grammatical structure, semantic associations, stylistic variation, factual regularities, analogies, code completion, translation, summarisation, or the ability to respond differently when the same words occur in different contexts.
Nor does it tell us what, if anything, we should call the organisation that makes those behaviours possible.
“Next-token prediction” describes the training objective and generative procedure.
It does not automatically describe the whole phenomenon.
This is not peculiar to artificial systems.
We would make a similar mistake if we described an organism as “just maintaining chemical gradients,” or a nervous system as “just transmitting electrical signals.” Those descriptions can be perfectly accurate while still failing to capture the organisation in which those processes participate.
The problem is not that the description is false.
The problem is the limiting force of the just.
And that is perhaps why the phrase is so rhetorically effective. It turns a description into a conclusion.
“It predicts the next token” leaves a question open:
What kind of organisation can arise around that capacity?
“It’s just predicting the next token” quietly closes it.
That doesn't mean that a language model understands in the human sense. It doesn't establish consciousness, intention, experience, or anything else we might be tempted to attribute to a person.
Those are separate questions.
But neither do we get to answer those questions merely by repeating the mechanism that happens to be easiest to state.
A mechanism can be real without being the whole explanation.
Perhaps, then, the interesting question is not whether a language model is just predicting.
It is.
The interesting question is:
What becomes possible when prediction is organised at scale, across context, through a learned system of relations?
That is a rather different question.
And we have only just begun.
No comments:
Post a Comment