"English is not learnable."
-- Noam Chomsky
Chomsky said this with his usual calm, impassive, casually smooth voice of confidence as if it were obvious to everyone, as if this could not possibly raise a quizzical brow. The audience must be thinking, "Wait -- We all speak English here. We must have learnt it. You learned it! What are you talking about? I'm lost."
His little four-word sentence presents the greatest challenge to empiricism perhaps in all of Western history. Right or wrong, it's not trivial.
What he meant was that language can't be learnt by mere empirical observation. Let me show you something remarkable about English. Take a sentence like
The drunk on the chair with three legs at the end of the bar wants to talk to you.
That sentence is mostly a sequence of prepositional phrases: a preposition followed by a noun phrase (an article followed by a noun). Here are the prepositions italicized, and the noun phrases underlined
The drunk on the chair with three legs at the end of the bar in the suit with the stripes wants to talk to you.
Notice that it's not the three legs that are at the end of the bar, it's the chair (with the three legs) that's at the end of the bar. The "three legs" stands between "the chair" and "the end", and yet English speakers can put these distant phrases together meaningfully. Now take a look at this shorter, simpler sentence
The drunk at the end in the suit of the bar
simpler yet impossible to connect "the end" to "of the bar". It's not harder to parse; it's impossible. It's not within the structure of English grammar.
If a learner can learn the longer sentence and learn to connect phrases at a distance with irrelevant information in between, why can't the learner accomplish the same task with the shorter one?
The structure of the longer sentence should teach the learner to parse the shorter one. But it doesn't. English speakers don't produce the shorter one, but do parse and produce the longer, more complex one, even though the parts are made of the same structural pieces.
This problem is called the "poverty of the stimulus": whoever acquires English must have learnt without evidence that the short sentence is ungrammatical while the longer more complex one is okay.
It requires some innate structural mechanism to produce innovative sentences while not producing others, and these twin abilities -- producing grammatical novelties while avoiding ungrammatical novelties made of the same parts as the grammatical ones -- cannot be learnt through observation, or even through observation of an absence.
Now, suppose the mind has a machine structure capable of churning out innovative sentences, and that machine is mechanically so structured that it cannot mechanically produce the non sentences, the way a touring bicycle's rear wheel can rotate forward and backward but the pedals can only engage the forward rotation not the backward because of the mechanical structure of the hub -- the structure allows the pedals to engage the wheel one way and not the other. In 1956, Chomsky identified a computational machine that could easily churn out those long, complicated sentences, but mechanically couldn't produce the simpler one, demonstrating that the brain must have such a structure. The mind has an innate and necessary contribution to language learning. It's comparable to Immanuel Kant's speculation that space, time and causality are innate human mental faculties that we contribute to our perception and understanding of the phenomenal world. Kant was answering the brilliant skeptical empiricism of Hume. Chomsky was answering the brilliant skeptical positivist behaviorism of Wittgenstein and Quine.
Tai-Danae Bradley has promoted the Yoneda lemma in category theory as a proof that neural networks are capable of capturing all phenomenal knowledge. The lemma entails that an object can be completely understood and described through all its possible relations to all other objects in the world to which it and they belong. This perspective on knowledge contrasts with the old Aristotelian account of objects analyzing the properties of the object itself, its material, its form, its purpose, its origin. Instead, the Yoneda perspective purports to either derive these properties -- the form and purpose, for example -- from the object's relations, or it dismisses them as irrelevant to the understanding of the object as it is. For example, the origin and the material of a word -- word tokens can have spoken form or ink or pencil or digital material tokens, and these differences do not change the meaning of the word. The Yoneda perspective is admirably suited to language.
The origin or cause of an object might seem essential to us in our Aristotelian and innately predictive mind -- natural selection has given us a drive to theorize predictively in order to protect us from dangers, and causality is essentially prediction of effects -- but one of Saussure's foundational linguistic insights was that the meaning and value of a sign is its relations at any particular moment. Its past has no effect on its current meaning. "Silly" once meant "blessed" in old English, but just try addressing New York's Cardinal as "Your silliness". How the word changed is an academic question, not a matter of current usage and meaning.
So setting aside Aristotle's effective cause and our innate protective, predictive penchant for explaining everything by its effective cause, the Yoneda perspective confronts a different problem of learning. It seems that the relations of prepositional phrases include long distance modification, yet the less long distance of the short sentence is excluded from the grammar. If we already know that the reason is a mechanical one, and we already know the machine that generates the complex structures and that cannot mechanically generate the shorter one, then of course we can affirm that the Yoneda lemma holds. But if we don't know already that hermetic information, the Yoneda perspective seems to fail.
Put differently, the Yoneda lemma works as a learning method only for objects that can be observed behaviorally, not for objects that belong to an unknown structure. Quine attempted to provide examples of such hermetic structures: the English notion of a rabbit compared with a culture that views the animal as a collection of functional parts. Behaviorally, viewers of the animal from both cultures will respond alike. How would the Yoneda perspective distinguish the two without cultural inside information.
In other words, the Yoneda lemma suffers from the weakness of empiricism and behaviorism.
You might say, so what and who cares? Well, it's not just an abstruse old philosophical boring debate. It's the heart of neural networks. The Shannon model of learning is essentially a behaviorist, empiricist model of learning. Chomsky's model of acquisition -- the Kantian mind-contribution view -- is a Turing model: there are some structures that require an innate computational machine. In the case of English prepositional phrases (these are not Chomsky's exemplars, but I choose them because they are more transparent and simpler and Chomsky's are beset with his own theoretical baggage), knowing the machine is necessary for learning the productive grammar and preventing the impossible ungrammatical strings.
So the problem of AI learning is a very old one. It goes all the way back to Plato and Aristotle, the medieval nominalist debate, to Hume and Kant's response to Hume, to the logical positivists and Wittgenstein, to Turing and Shannon, and now Chomsky and the AI engineers and Tai-Danae Bradley.
One final point here. The goal of engineering is to accomplish a task for some consumer, whether market consumer, military or government consumer. How the task is achieved is of little importance as long as its benefits are greater than its costs. If AI can generate English prepositions and reject the impossible ones, it has succeeded in learning English. If it can also learn impossible structures or impossible languages that humans can't, well that's great for the AI engineers but it tells us that whatever machine structure the AI is using it's not what humans are using. In other words, the fact that AI can learn English doesn't tell us anything about how we humans acquire English.
Now, humans can learn beyond their innate language faculty, the way humans learn reading, writing and 'rithmatic -- through lots of repetitive unnatural work. This kind of learning compared with first language learning is analogous to learning to ride a bike compared with learning to walk. One is hard to learn, the other hard not to learn. So the question to ask of LLMs is, do they learn everything with equal ease? Neither LLMs nor their engineers know the answer...yet.
As it happens, AI can't even introspect to find out and explain how it learns English. AI has no more introspective ability than humans have -- not only do most people not know how they learnt their native language, they often, maybe mostly, find the explanations they're given in grade school difficult to understand or recognize, and that's even so of linguists who spend lots of grant money on trying to explain in detail how we learn. AI's "reasoning" is just noting its step-by-step response to a prompt, not explaining how it learnt to do that (the Myth of Introspection: AI introspection, like human introspection (bullshit)). If you ask it how it learnt, its explanation is more like a search function. It's not introspection into how it learnt; it's a search through the literature to see what the literature says about how it learned. It's a bit of a comedy. One linguist described LLMs as autocoprophagy -- auto- (self) copro- (shit) phagy (eating) -- LLMs are sort of us eating our own shit and shitting it out. I think the linguist intentionally added an "r" to the "copro" as "autocorprophagy" to indicate the driving role of the AI corporations. Well done, covered all bases.
For the linguist, this hard learning presents an obstructive confound for Chomsky's theory. Remember that the key characteristic of the computational machine is its ability to generate new sentences following, and never violating, the machine structure's generative capacity. But if humans have a general capacity to hard-learn generative language strings, how does the linguist know which innovative sentence strings are evidence of the machine structure versus the general hard-learning capacity? Once a child learns to ride a bike, the child can ride all sorts of bikes -- mountain, stunt, folding, fix-wheel, multiple-speed. That's a productive capacity generated by general hard learning. In the case of bike-riding, we can distinguish the hard-learning from the innate learning of walking: we observe toddlers learning to walk by themselves and observe that learning to ride a bike needs a lot of push. With language, the hard learning happens internally, so it's not clear that we are able to distinguish the evidence of the language faculty from evidence of general learning. And if the linguist can't distinguish which evidence is evidence of the language faculty and evidence of hard general learning, all the evidence of the language faculty may as well be in a black box. If the bottle of olive oil on the store shelf isn't labeled "Moroccan", how do you know it's not from Greece without opening and tasting it?
The conclusion has to be that even if the Chomsky model of language learning is completely accurate, the program of discovering its details cannot succeed, given the current inaccessibility of the evidence. This was my conclusion in the 90's when I was getting my PhD in linguistics, and one reason I didn't devote myself to syntax. (Also semantics, which is a relation between language and information and understanding of the world, is just more interesting to me than grammatical structure.)
Summing up, the goals of technology are simpler than those of science. The details of how the technology succeeds don't matter to the engineer or to the technology as long as it succeeds. Engineers don't need to know whether a sentence structure is hard-learnt or innately machine-generated. It doesn't matter if a sentence structure is a historical vestige that is merely lingering because it was so frequent in the language that it has survived. It also doesn't matter if it is a borrowing from another language's structure. If the technology is successful, it'll learn it all equally well.
The superficiality of technology and engineering is one of its strengths. The inventors of the bicycle didn't know that in order for the bike to turn, the rider has to learn to one side, otherwise the bike will fall over. The engineer didn't need to know that. This is common among technologies. Science, however, has a much higher bar and a more difficult task. Science must explain. And the complexity of the world presents constant confounds to the scientific theory -- dark matter, dark energy -- that mustn't be ignored. The depth of science is its weakness.
In short, LLMs prove that language can be learnt. But it doesn't tell us how.
If this was interesting, take a look at four AI myths of human intelligence.