There is another phrase that frequently appears in discussions of large language models:
“It’s only statistics.”
It sounds reassuring.
Statistics, after all, is what happens when you count things.
And counting things seems a long way from thinking.
A language model, on this account, has no understanding of language. It has no concepts, no intentions, no meanings. It has simply absorbed statistical regularities from vast quantities of text and learned to reproduce them.
There is something important here.
Language models really do depend upon statistical regularities.
But then comes the word only.
Like just, it does not add to the description of what the system does. It limits the significance we are permitted to attach to it.
And once again, the limitation is doing more work than the explanation.
Consider what statistics can contain.
The relationship between two variables can be statistical because there is some underlying regularity connecting them. A population can exhibit statistical structure that no individual member possesses. A statistical distribution can encode the accumulated consequences of countless interactions. A correlation can reveal something about a system even when no single event explains it.
Statistics is not necessarily the opposite of structure.
Sometimes statistics is how structure becomes visible.
Language is an especially obvious case.
No speaker has memorised every sentence that can be said in English. No dictionary contains all the possible relationships among words. Yet linguistic behaviour is not random.
Words occur in patterns.
Sounds occur in patterns.
Grammatical structures occur in patterns.
Meanings constrain one another.
Genres develop expectations.
Conversations establish local probabilities.
A sentence changes what can naturally come next.
The statistical regularities of language are therefore not simply a heap of numbers sitting behind language. They arise from an enormous history of people using language in social contexts.
A model trained on that history is consequently doing something rather peculiar.
It is not merely counting words.
It is acquiring a system whose parameters have been reorganised by exposure to the statistical structure of language.
That distinction matters.
A dictionary contains information about words.
A corpus contains patterns of use.
A language model learns a high-dimensional organisation of those patterns.
We can describe that organisation statistically because statistics is the appropriate mathematics for describing it.
But the mathematical description does not necessarily exhaust the organisation being described.
This is a familiar problem elsewhere.
A temperature is a statistical property of a physical system. That does not make temperature unreal because no individual molecule has the temperature of the room.
Population genetics is statistical. That does not mean that organisms are imaginary.
Evolutionary change can be described statistically. That does not mean that nothing happens to individual organisms.
The fact that something can be described statistically tells us something about the form of the description.
It does not, by itself, tell us what kind of thing is being described.
And there is an especially interesting irony here.
The phrase “only statistics” often seems to imply that statistical processes cannot produce anything genuinely structured.
But much of science consists precisely in discovering that large numbers of local interactions can produce stable global organisation.
A gas has pressure.
A population has a distribution.
A market has patterns.
A flock has a form.
An ecosystem has dynamics.
None of these requires that every component possess a representation of the whole.
The organisation can exist at a different level from the individual interactions that participate in producing it.
This does not mean that a language model is therefore equivalent to an organism, a society, or a mind.
It isn't.
The differences are substantial, and precisely what those differences imply remains an empirical and conceptual question.
But “only statistics” cannot answer that question.
It tells us that the system's behaviour is statistically organised.
It does not tell us what that organisation amounts to.
And perhaps there is an even more uncomfortable point.
Human language itself is statistical.
Not merely because computers can count it.
Human speakers have expectations about what can come next. We recognise familiar patterns. We exploit probabilities. We become surprised when an unlikely expression occurs. We adapt our linguistic behaviour to context.
Much of what we call linguistic competence consists in being extraordinarily good at navigating a structured field of possibilities.
So if statistical regularity disqualified a system from possessing anything interesting, we would first have to explain why statistical regularity does not disqualify us.
Of course, the human case is not the machine case.
Human linguistic behaviour is embedded in bodies, histories, relationships, purposes and lived environments.
That difference matters enormously.
But notice what follows.
It gives us a reason to investigate the organisation in which statistical regularity participates.
It does not give us a reason to dismiss statistics itself.
Perhaps, then, “It’s only statistics” has the same problem as “It’s just prediction.”
The statement begins as a description.
Then only turns it into a dismissal.
But statistical organisation is not the absence of organisation.
It is one way organisation can exist.
And once we recognise that, a rather more interesting question appears:
What kind of organisation can statistics become when a system learns from the statistical structure of an entire linguistic world?
That is a question worth asking.
“Only statistics” is not an answer.
No comments:
Post a Comment