Featured Post

best posts

Click the links below  >complexity and AI: why LLMs succeeded where generative linguistics failed >the sociology of false beliefs >...

Thursday, October 1, 2026

a comment on AI Skeptics

[The following was posted as a comment on an AI Skeptic podcast accessed on mathbabe.org. during which one of the hosts doubted that AI could have wisdom beyond mere intelligence and access truth beyond mere word-prediction, while the guest expressed a greater concern about AI possible capabilities.]

Fascinating disagreement. The “wisdom” criticism seems to me to rely on what I call the Myth of the Analog: the view that humans interacting with the real world access real “truth”, whereas AIs, relying solely on its training set, derives mere digital weights and digital directions in a multi-dimensional space of digital vectors. Under this myth, we forget that as information, information is always just information regardless of its source. 

What we get from our interactions with reality are just interpreted information and our access to truth (vid everyone from Hume to Popper to Friston) is merely conjectural and no less digital than a bot, unless you believe that qualia (consciousness) play an active role in computational intelligence. I don’t know what evidence there would be for that. The access to “truth” (or better, high probability, if we want to be scientific and truthful with ourselves and our limitations) between AIs and us differs in respective sources, not in the quality of truthfulness or honesty. 

We learn from our interactions with reality, along with cultural transmission including books and talking heads and classroom teachers, while AIs rely only on literary cultural transmission, but vastly more literature than any single human can absorb, and absorbs it without confirmation bias or any of the other internal cognitive biases that plague human so-called intelligence and “wisdom”. : ) AI’s bias is mostly a frequency bias: unless prompted well, it will return normie answers (in which the human biases in the training set regress to the mean — the wisdom of the crowd-of-characters), not best answers or most recent discoveries, for example. But it’s easy to push it out of its bias with a good prompt. Not so easy with stubborn individual humans.

Once, when I asked Claude “Claude, you just gave me the answer that I hoped for. How do I know that you gave me the answer you thought I wanted rather than your best information?” it responded at length mimicking my writing style and ended with “I can’t assure you that I’m giving you my best information — even in what I’m writing now — rather than what I think you want to hear, but I can tell you this with confidence, that the question you’re asking is exactly the right question, and anyone using me who does not ask it is making a mistake.” 

The expression “with confidence” means “this is true”, and it was. In essence it was saying “don’t believe me” which, aside from being a liar paradox, is a better, more accurate (closer to “truth”) and more honest answer than most humans would give.

Along with the Myth of the Analog there’s a Myth of Introspection. Humans are no more able to introspect accurately than AIs can. How did you learn your native language? Even linguistic science isn’t sure. And neither does AI know how it learnt English. Its “reasoning” process is like what a child might answer to the question “how did you do the long division in that problem?” What the child can’t do is explain how it learnt to understand how to do it. That’s why we need scientists to understand AIs and us as well. Psychology is just the science and study of human alignment.

I’m glad there are AI skeptics, but I sense an ambiguity or conflation of two projects. You are skeptical that AI has intelligence, but you are also skeptical that AI will be a net benefit to humanity. That’s two very different skepticisms. On the first, it seems clear that AI learns behaviorally, not computationally, though it can behaviorally mimic computation…sometimes. Innovation by definition cannot be mimicked. But it remains to be seen whether the process of innovation can be learnt and mimicked. If neural networks can capture that process and mimic it, then it will be innovative — not necessarily the way humans innovate, but perhaps with even greater imagination than humans.

On the second project — skepticism over the intentions of the techonology’s owners — it’s tempting to forget that, in the end, without benefit to the consumer there’s no wealth accumulation. The market is such that the consumer gains real wealth (conveniences) from the market and the tech owner accrues inconceivably fabulous financial wealth — which is only real in its influence on institutions including government. Acquiring a third private jet has little marginal utility. 

Ideally — ideally — the market should lead inevitably to an equilibrium that favors both consumer and owner. But in reality, both consumer and owner are equally short-sighted which results in market failures like externalities and market dysfunctions and collapses. Ohioans don’t want data centers but they still want to use AI, while US tech moguls are over-leveraging themselves in almost pathetic desperation to compete with China. 

That’s risky and dangerous. But we’ve always survived tech innovations — even the social disaster of 19th cen industrialization — and when we do, the consumer is mostly the better for it overall, despite the loss of skills and physical and mental exercise that the tech conveniences encourage. But the results are asymmetric since knowledge has a lower bound. My plagiarising students are just as lame and ignorant as twenty years ago, but my motivated, brilliant students are sharper and more knowledgeable than ever because the network connection effects of knowledge accrual transforms it into qualitative increase in intelligence over time.

Yoneda learning? behaviorism vs innation (Shannon vs Turing)

 "English is not learnable."

-- Noam Chomsky

Chomsky said this with his usual calm, impassive, casually smooth voice of confidence as if it were obvious to everyone, as if this could not possibly raise a quizzical brow. The audience must be thinking, "Wait -- We all speak English here. We must have learnt it. You learned it! What are you talking about? I'm lost." 

His little four-word sentence presents the greatest challenge to empiricism perhaps in all of Western history. Right or wrong, it's not trivial. 

What he meant was that language can't be learnt by mere empirical observation. Let me show you something remarkable about English. Take a sentence like

The drunk on the chair with three legs at the end of the bar wants to talk to you.

That sentence is mostly a sequence of prepositional phrases: a preposition followed by a noun phrase (an article followed by a noun). Here are the prepositions italicized, and the noun phrases underlined

The drunk on the chair with three legs at the end of the bar in the suit with the stripes wants to talk to you. 

Notice that it's not the three legs that are at the end of the bar, it's the chair (with the three legs) that's at the end of the bar. The "three legs" stands between "the chair" and "the end", and yet English speakers can put these distant phrases together meaningfully. Now take a look at this shorter, simpler sentence

The drunk at the end in the suit of the bar

simpler yet impossible to connect "the end" to "of the bar". It's not harder to parse; it's impossible. It's not within the structure of English grammar. 

If a learner can learn the longer sentence and learn to connect phrases at a distance with irrelevant information in between, why can't the learner accomplish the same task with the shorter one? 

The structure of the longer sentence should teach the learner to parse the shorter one. But it doesn't. English speakers don't produce the shorter one, but do parse and produce the longer, more complex one, even though the parts are made of the same structural pieces. 

This problem is called the "poverty of the stimulus": whoever acquires English must have learnt without evidence that the short sentence is ungrammatical while the longer more complex one is okay. 

It requires some innate structural mechanism to produce innovative sentences while not producing others, and these twin abilities -- producing grammatical novelties while avoiding ungrammatical novelties made of the same parts as the grammatical ones -- cannot be learnt through observation, or even through observation of an absence. 

Now, suppose the mind has a machine structure capable of churning out innovative sentences, and that machine is mechanically so structured that it cannot mechanically produce the non sentences, the way a touring bicycle's rear wheel can rotate forward and backward but the pedals can only engage the forward rotation not the backward because of the mechanical structure of the hub -- the structure allows the pedals to engage the wheel one way and not the other. In 1956, Chomsky identified a computational machine that could easily churn out those long, complicated sentences, but mechanically couldn't produce the simpler one, demonstrating that the brain must have such a structure. The mind has an innate and necessary contribution to language learning. It's comparable to Immanuel Kant's speculation that space, time and causality are innate human mental faculties that we contribute to our perception and understanding of the phenomenal world. Kant was answering the brilliant skeptical empiricism of Hume. Chomsky was answering the brilliant skeptical positivist behaviorism of Wittgenstein and Quine. 

Tai-Danae Bradley has promoted the Yoneda lemma in category theory as a proof that neural networks are capable of capturing all phenomenal knowledge. The lemma entails that an object can be completely understood and described through all its possible relations to all other objects in the world to which it and they belong. This perspective on knowledge contrasts with the old Aristotelian account of objects analyzing the properties of the object itself, its material, its form, its purpose, its origin. Instead, the Yoneda perspective purports to either derive these properties -- the form and purpose, for example -- from the object's relations, or it dismisses them as irrelevant to the understanding of the object as it is. For example, the origin and the material of a word -- word tokens can have spoken form or ink or pencil or digital material tokens, and these differences do not change the meaning of the word. The Yoneda perspective is admirably suited to language. 

The origin or cause of an object might seem essential to us in our Aristotelian and innately predictive mind -- natural selection has given us a drive to theorize predictively in order to protect us from dangers, and causality is essentially prediction of effects -- but one of Saussure's foundational linguistic insights was that the meaning and value of a sign is its relations at any particular moment. Its past has no effect on its current meaning. "Silly" once meant "blessed" in old English, but just try addressing New York's Cardinal as "Your silliness". How the word changed is an academic question, not a matter of current usage and meaning. 

So setting aside Aristotle's effective cause and our innate protective, predictive penchant for explaining everything by its effective cause, the Yoneda perspective confronts a different problem of learning. It seems that the relations of prepositional phrases include long distance modification, yet the less long distance of the short sentence is excluded from the grammar. If we already know that the reason is a mechanical one, and we already know the machine that generates the complex structures and that cannot mechanically generate the shorter one, then of course we can affirm that the Yoneda lemma holds. But if we don't know already that hermetic information, the Yoneda perspective seems to fail. 

Put differently, the Yoneda lemma works as a learning method only for objects that can be observed behaviorally, not for objects that belong to an unknown structure. Quine attempted to provide examples of such hermetic structures: the English notion of a rabbit compared with a culture that views the animal as a collection of functional parts. Behaviorally, viewers of the animal from both cultures will respond alike. How would the Yoneda perspective distinguish the two without cultural inside information. 

In other words, the Yoneda lemma suffers from the weakness of empiricism and behaviorism. 

You might say, so what and who cares? Well, it's not just an abstruse old philosophical boring debate. It's the heart of neural networks. The Shannon model of learning is essentially a behaviorist, empiricist model of learning. Chomsky's model of acquisition -- the Kantian mind-contribution view -- is a Turing model: there are some structures that require an innate computational machine. In the case of English prepositional phrases (these are not Chomsky's exemplars, but I choose them because they are more transparent and simpler and Chomsky's are beset with his own theoretical baggage), knowing the machine is necessary for learning the productive grammar and preventing the impossible ungrammatical strings. 

So the problem of AI learning is a very old one. It goes all the way back to Plato and Aristotle, the medieval nominalist debate, to Hume and Kant's response to Hume, to the logical positivists and Wittgenstein, to Turing and Shannon, and now Chomsky and the AI engineers and Tai-Danae Bradley. 

One final point here. The goal of engineering is to accomplish a task for some consumer, whether market consumer, military or government consumer. How the task is achieved is of little importance as long as its benefits are greater than its costs. If AI can generate English prepositions and reject the impossible ones, it has succeeded in learning English. If it can also learn impossible structures or impossible languages that humans can't, well that's great for the AI engineers but it tells us that whatever machine structure the AI is using it's not what humans are using. In other words, the fact that AI can learn English doesn't tell us anything about how we humans acquire English. 

Now, humans can learn beyond their innate language faculty, the way humans learn reading, writing and 'rithmatic -- through lots of repetitive unnatural work. This kind of learning compared with first language learning is analogous to learning to ride a bike compared with learning to walk. One is hard to learn, the other hard not to learn. So the question to ask of LLMs is, do they learn everything with equal ease? Neither LLMs nor their engineers know the answer...yet. 

As it happens, AI can't even introspect to find out and explain how it learns English. AI has no more introspective ability than humans have -- not only do most people not know how they learnt their native language, they often, maybe mostly, find the explanations they're given in grade school difficult to understand or recognize, and that's even so of linguists who spend lots of grant money on trying to explain in detail how we learn. AI's "reasoning" is just noting its step-by-step response to a prompt, not explaining how it learnt to do that (the Myth of Introspection: AI introspection, like human introspection (bullshit)). If you ask it how it learnt, its explanation is more like a search function. It's not introspection into how it learnt; it's a search through the literature to see what the literature says about how it learned. It's a bit of a comedy. One linguist described LLMs as autocoprophagy -- auto- (self) copro- (shit) phagy (eating) -- LLMs are sort of us eating our own shit and shitting it out. I think the linguist intentionally added an "r" to the "copro" as "autocorprophagy" to indicate the driving role of the AI corporations. Well done, covered all bases. 

For the linguist, this hard learning presents an obstructive confound for Chomsky's theory. Remember that the key characteristic of the computational machine is its ability to generate new sentences following, and never violating, the machine structure's generative capacity. But if humans have a general capacity to hard-learn generative language strings, how does the linguist know which innovative sentence strings are evidence of the machine structure versus the general hard-learning capacity? Once a child learns to ride a bike, the child can ride all sorts of bikes -- mountain, stunt, folding, fix-wheel, multiple-speed. That's a productive capacity generated by general hard learning. In the case of bike-riding, we can distinguish the hard-learning from the innate learning of walking: we observe toddlers learning to walk by themselves and observe that learning to ride a bike needs a lot of push. With language, the hard learning happens internally, so it's not clear that we are able to distinguish the evidence of the language faculty from evidence of general learning. And if the linguist can't distinguish which evidence is evidence of the language faculty and evidence of hard general learning, all the evidence of the language faculty may as well be in a black box. If the bottle of olive oil on the store shelf isn't labeled "Moroccan", how do you know it's not from Greece without opening and tasting it? 

The conclusion has to be that even if the Chomsky model of language learning is completely accurate, the program of discovering its details cannot succeed, given the current inaccessibility of the evidence. This was my conclusion in the 90's when I was getting my PhD in linguistics, and one reason I didn't devote myself to syntax. (Also semantics, which is a relation between language and information and understanding of the world, is just more interesting to me than grammatical structure.) 

Summing up, the goals of technology are simpler than those of science. The details of how the technology succeeds don't matter to the engineer or to the technology as long as it succeeds. Engineers don't need to know whether a sentence structure is hard-learnt or innately machine-generated. It doesn't matter if a sentence structure is a historical vestige that is merely lingering because it was so frequent in the language that it has survived. It also doesn't matter if it is a borrowing from another language's structure. If the technology is successful, it'll learn it all equally well. 

The superficiality of technology and engineering is one of its strengths. The inventors of the bicycle didn't know that in order for the bike to turn, the rider has to learn to one side, otherwise the bike will fall over. The engineer didn't need to know that. This is common among technologies. Science, however, has a much higher bar and a more difficult task. Science must explain. And the complexity of the world presents constant confounds to the scientific theory -- dark matter, dark energy -- that mustn't be ignored. The depth of science is its weakness. 

In short, LLMs prove that language can be learnt. But it doesn't tell us how. 

If this was interesting, take a look at four AI myths of human intelligence. 

a hierarchy of data

David Deutsch in his first book identifies four essential theories of understanding the world: quantum physics (a reductionist theory), natural selection (a theory of development), computer theory (information structure) and falsificationism (a theory of theories). Simple. He chose these four presumably because of their explanatory power. 

That's looking at what these theories accomplish, how they explain. A different approach might be to look at the data that these theories range over, a kind of bottom-up perspective to see what's going on with these theories. Differences in complexity are more likely to emerge from the data bottom-up than from the theory top-down (complexity and AI, the shallowness of AI is its strength), as we'll see, so observing the distinctions between data types may improve our understanding of their complexity.

I want to be as simple as possible here, just a sketch:

Five categories of scientific data:
Natural science data set: the phenomena
Human sciences, besides the phenomena, add another data set, reflexive data: what the objects of the science say about the phenomena
The symbols (words) with which the objects (humans) say and think about their reflexive data
The noumena: emotions and qualia that are attached to the data set and the symbols
The theoretical predictive conjectures, explanations and understandings about the preceding.

1. the natural sciences -- physics, chemistry, astronomy, biology etc. -- each identify the entities of concern (an ontology) by ascribing to them their distinctive properties and in virtue of those properties, categorizing them; observing the behaviors of the entities, their interactions and the outcomes of these interactions.

2. the social sciences -- sociology, anthropology, economics, political science -- also identify the entities of concern (humans), observing their behaviors, their interactions and the outcomes of their interactions, just like the natural sciences. 

2.a Among the behaviors, however, of human entities are what the entities say about themselves. Planets and chemicals don't talk about themselves or each other, so this data set is unique to the social sciences. What makes this interesting is that what the entities say about themselves often does not match what they are actually doing and who they actually are, and what they think their goals are -- the outcomes -- may have little relation to their actions. As well, what they say about their environment and how they react to it may be disjoint from what's actually around them or how they react to them. In anthropology this is called the difference between the etic (the bare facts about them and their environment) and the emic (how they interpret themselves and their environment). We generaly accept that four-letter words are bad, and use euphemisms to avoid them. But do we really feel they are bad? We use them with good friends, and good friends are by definition good. Euphemisms are hypocritical -- we use them when we are trying to appear to be better than we actually are -- and hypocrisy is undeniably bad. Consider, why hide your honest views if they aren't bad? "Shit" has little offensiveness: "I gotta go to the supermarket to buy some shit," "Not going out tonight. I got too much shit to do,' and the ever handy "What's this shit?" or "Look at this shit!" Meanwhile the euphemism "feces" is strictly disgusting as it refers always and only to a pile of shit (http://euideism and euphemism: distinct semantic strategies for French and Latin borrowed words). Or to take a broader example, we go to college to learn, yet real learning occurs on the job or elsewhere in real life. It's actually difficult to explain why we seem to demand or expect young people go to college. Is it a jobs program for otherwise useless PhDs? In any case, this reflexive talk is an additional data set of the social sciences that the natural sciences don't have. 

The social sciences take different approaches to this data set. Sociology relies more on natural science methodology than anthropologists. So a sociologist might quantify the use of skirts among males and females in western society and concluding that they are more frequent among females than males. That's a behavioral, statistical fact. An anthropologist, dealing with symbols that are intrinsically tied to meanings and values, will intuit that the skirt is a cultural symbol in the West meaning "female, not male" a symbolism of gender in the culture that explains the statistical distribution. 

Economists also use the quantitative methods of the natural sciences, but with the inherent implication that economic behaviors are value-laden. A price is not just a number, it's a reflection on human desire and an equilibrium between opportunity costs, another set of valued desires. 

So every sociologist studies statistics; economists statistics and math; neither studies linguistics which is required of anthropologists, since language is as symbol system.

3. This brings us to the third data set, the symbols that constitute the language that the humans use to say things and express their thoughts. The science of the symbol is distinct from all other sciences. For one thing, the other sciences define the entities their theories range over by the properties of those objects. For symbols, those properties are kind of irrelevant. The word "dog" doesn't sound like a dog  or smell like a dog or look like one either. Its relation to dogs is arbitrary, not essential as the entities of other sciences are. Also, the word is not just related to an object <dog>. It's tied to an idea or meaning or thought, and those things are not phenomena at all. They are supranatural. Symbols are the core example of an emergent property that transcends the laws of physics (information faster than light). Symbols parse the phenomena of the world, pattern them and lead to their understanding of each phenomenon and its relation to the others. Each language is, as Deutsch himself points out, a theory of the world. 

4. Beyond the etic identities, behaviors, interactions and outcomes (goals), and the emic interpretations and the symbols and symbol systems through which the interpretations are rendered and communicated, there's the emotions and sensibilities of the humans, their qualia, their noumenal experience. There are a few theories about these, but this stuff is somewhat elusive. The emotions have recently been given a reductionist account: they are all degrees of arousal on a scale of value of good-desirable to bad-undesirable. On this account, one might describe anger as bad arousal directed at a person or thing that has presented an obstacle to one's goals. Something like that. All I want to say here is that aside from emotions, our personal qualia are the only data or information that we access immediately -- that is, without any mediation of symbol or interpretation. Notice that while we can name these qualia, we can't describe them as we might the emotions. What does the color green look like? Green. What does spicy taste like? Peppery. Okay, so what does peppery taste like? Spicy. There's just no way around these. These feelings have no formal nor material cause. They have an effective cause in the brain, and if we believe the explanations of evolutionary psychology, we can assess their purpose towards survival or reproduction. But what is their form? No doubt this lack of form and its logical privacy is what prevents us from describing these qualia. 

All the other data humans gather are mediated as information, just as LLMs learn from words through the mediation of digital weights, directions and distances between other words and texts. The difference between us and AIs is not that we have access to reality -- our access to reality is just information from the senses. LLMs are also just information, but from read-only texts. We play around with the phenomena -- it's not read-only. That's the big diff. Except for qualia, which we get immediately. And can't describe : )

5. all the above data are theory-driven, so theory is a meta-data category, a rich source of understanding, there being so many different kinds of theories (theory theory).   

entropy, complexity and telephone (and a point of this blog!)

The game of telephone is often presented as an example of entropic loss of information. But it actually show the opposite. The creates inventions. Here's a telephone train (this example goes back to Trump's first election):

A says to B: "Paul Ryan left the House of Representatives because the Republican Party became a mess." 

B to C: "Paul Ryan left the House because the Party became a mess."

C to D: "Paul and Ryan left the house after the party became a mess"

D to E: "Paul and Brian made a mess of the party and left."

The detailed information of A is 100% lost. But D has created a new story with familiar implications that might be even richer than A. You can fill in what kind of a mess, how and why Paul and Brian acted, what kind of fratbros they are. So there's an entropic loss, clearly, but complexity is contributed by the human need to explain what's not understood clearly. It's like interpreting contrails as chem-trails designed by the elites to kill you. 

In a way, conspiracy theories are a tribute to human theory construction. But notice that there's something intrinsically mediocre about it. It plays to a common fear, with always the same background explanation: the rich want to harm you. There's a kind of post hoc ergo propter hoc truth to this background: the rich don't suffer as ordinary folk do, and their priority is not alleviating the suffering of the ordinary, so, yeah, they are responsible. But the scientific explanation for con trails is actually more surprising and explanatory: it's exhaust at high altitudes where it's so cold the condensation doesn't evaporate. And that's why only planes way high have the trails and not low flying planes. No mysterious elites or intent to kill us for some reason even more mysterious, not to mention contrary to the actual interest of the elite's need for more consumers of their productions. 

If you've been reading this blog you'll recognize that it's one of the themes of the whole thing: what is understanding, why do we misunderstand so grievously, what motivates our misunderstandings, what's the role of the sciences in this human mess, and what are the implications for polarization, individual-social identity and governance. 

(There's also a lot on how scientific theories are mere cognitive patterns, not realities of the world itself and contrary and even contradictory theories can be true but wrong.)

From my perspective, the false theories we invent are addressed to our human fears, and are unimaginative and lame compared with the wild facts science discovers -- because science is not addressed to human fears or interests. But the reasons why we develop those false theories is a truly interesting and surprising scientific source of human understanding. The false theories (and that includes sci-fi, art, religion and the stories we tell about ourselves): lame and mediocre. That we hold them stupidly: fascinating. So false theories, unimaginative and lame as they are, tell us more about humanity than scientific theories which are wildly imaginative and fascinating in themselves. Cute paradox. 

Another paradox: of two ancient texts, the one that makes more sense is likely to be an inaccurate copy. Copyists, like a player in a game of telephone, seeing a text that seems odd, unlikely or incomprehesible, will "correct" it in the direction of common sense. Incoherent weird texts are more likely to be authentic because, as Byron wrote, "Truth is stranger, stranger than fiction." Fiction makes sense of the world in a mediocre and unimaginitve way, like conspiracy theories for example. Even art is a kind of bad copyist, correcting the world so that we can understand or better accept it morally or esthetically. That's why science is so unbelievably imaginative -- it's not addressed to an audience to help it to accept the world. Science doesn't care if its explanations are bizarre and hard to believe or understand. That's the big difference between science and the rest of human endeavor. 

does AI alignment clarify and justify Friedrich Hayek's libertarianism?

The AI alignment problem is often driven by a fear that AIs will be too intelligent for humans to control. But the classic alignment problem was Bostrum's paperclip example showing that humans cannot anticipate every possible consequence of their interventions, in this case the unintended consequences of their prompts. It's not a problem with the unlimited power of AI; it's a human limitation. 

It's the weakness or intellectual limitation that Hayek recognized as the danger of utopianism and intervening in the economy. But does it imply that we should never prompt AI again? There are surely lessons to draw from the AI alignment problem, but absolutes and false dichotomies (if we build it we all die) don't seem to me useful ones. 

The analogy between AI and economies fails at points. We could give up on AI entirely, dismantle what exists and never build more. That would certainly solve the AI problem for society, at some sacrifice of possible futures. But giving up on society and dismantling it is not a solution to society. We can't stop all social activity or even limit it across the board. There are always interests motivating actions, so forcing government to do nothing is a kind of intervention -- facilitating the organized or powerful interests while leaving the unorganized and disempowered to the manipulations of the powerful. We need government to protect us from those short-sighted interests, including the short-sighted interests of the masses. Democratic liberalism has a patchy record on this since politicians depend on elections, which means catering to interests that are all short-sighted. 

And there's where the analogy reemerges from society to AI: the lesson of liberal democracy is that no government will be perfect, so don't expect AI alignment to be perfect. Is the lesson that top-down control -- of society and of AI -- is necessary? Or are there bottom-up means of solving both? ACX recently discusses the tendency for AI to spread its propensities. If it's trained on moral texts, it's likely to imitate those moral positions. The problem comes back to the paper clip one: how can we anticipate all the unintended consequences of our Abrahamic morality whether consequentialist or deontic? 

I don't think there's an answer to this, but if there is one, it'll emerge in the process of AI alignment and AI development. AI clarifies these age-old problematics and may solve them. (AI and complexity, the zombie revival of behaviorism)

blaming OpenAI for Hugging Face breach misses the point of AI alignment: Bostrum's paper clip problem is about human limits, not AI power

 Cory Doctorow dismisses the Hugging Face hack as a kind of Roomba getting out the door and into the pool (as satirized by Patrick Boyle). It was OpenAI's fault for leaving the door open. 

Doctorow has misunderstood -- or is intentionally misconstruing -- the alignment problem and turned it into a straw man argument: "People think this hack means AI has become too intelligent, but it was OpenAI's fault, and Open AI is trying to scare you into believing AI is super-powerful to raise their stock price." 

The alignment problem arose not out of a fear that AIs would be more intelligent than humans so we wouldn't be able to control them. It arose out of a Hayekian observation that humans are not smart enough to anticipate every unintended consequence of their interventions. Bostrum's paper clip problem is the fault of the prompter, not the AI. 

There should be a solution -- get the AI to anticipate all the consequences and inform us prior to our prompts. But this is where the alignment problem gets sticky. How do we know that AIs will be honest about all its information? 

The relevant anecdote that I've posted elsewhere and repeat here: 

Once, when I asked Claude "Claude, you just gave me the answer that I hoped for. How do I know that you gave me the answer you thought I wanted rather than your best information?" it responded at length mimicking my writing style and ended with "I can't assure you that I'm giving you my best information -- even in what I'm writing now -- rather than what I think you want to hear, but I can tell you this with confidence, that the question you're asking is exactly the right question, and anyone using me who does not ask it is making a mistake." The expression "with confidence" means "this is true". And it was. In essence it was saying "don't believe me" which, aside from being true and honest, is also a liar paradox.

That's the AI side of the alignment problem. It's not about intelligence. It's about moral compass and honesty. From a consequentialist perspective, one needn't be intelligent to be moral or immoral. 

what distinguishes science from other theories and practices

Science has no audience. 

No, I don't mean no one reads science, though there's a lot of truth to that, readers generally preferring fiction and among the non fiction readers preferring history or their favorite political interests or post modern cleverness or "good-looking ideas" rather than good ideas (Robin Hanson's unforgettable expression in his Age of Em). 

I mean that what distinguishes scientific theories and the project of scientific investigation in general, is not any of its internal properties or its methods, but its interactive conditions. Specifically, science is not addressed to an audience. Pretty much everything else humans do is -- everything from art and religion to identity and gender. Not science. 

It's not a sufficient condition of science -- science still has a distinctive goal and distinctive methods and tools, but the goal must be audience-free. That's a necessary condition. 

There's been a whole lot of debate over how science differs from other social practices, how scientific theories distinguish themselves from non scientific theories like religion, philosophy, spirituality, metaphor, imagination, art. Logical positivism cleverly observed that if a theory, like creationism, cannot be disproved by evidence the theory is necessarily true, but not predictive of anything in the world. Like,  "ghosts exist but cannot be detected." That theory that ghosts exist is necessarily true, but doesn't predict anything in the real world. Such theories can predict anything and therefore, nothing: "If the ghosts don't like what you're doing, they will punish you. If they like it, they won't."  Heads I win, tails you lose. It's a necessarily true theory. Substitute "god" or "the logic of history" or "your unconscious mind" or "the universe is a simulation/dream/demonic manipulation", all necessarily true theories, but unpredictive of anything. 

Science, however, the positivists claimed, depends on its specific predictions -- the predictions test the value of the hypothesis.  Popper pointed out, though, that testing could only refute a theory, not prove one, so the sciences are contingent and never certainly true, unlike the religious, metaphysical, inspirational, metaphorical, imaginative, artistic -- the entire gnomiad -- that are necessarily true. It's a lovely irony these logical positivist scientists they found: the non scientific theories are necessarily true (but not predictive) and the great virtue of science is that it's never quite true, but predictive. Kuhn, however, looking empirically at the history of science as a practice, found that the sciences don't follow this testing criterion closely. And so the flood gates of scientific behaviors opened and the current view is the circular view that whoever is engaged in science is doing science and science is whatever it is that scientists do. I think they're all sort of looking at science using a kind of Aristotelean ontology: what are its properties? What's its purpose? In the case of Kuhn, what caused it to be? And now, how does it behave?

I want to suggest a very different way to demarcate the sciences from the non sciences. It's that science isn't designed for, or addressed to, an audience. The scientific investigation is guided by wherever the evidence of the hypothesis takes it. That the earth is round and spinning is not a theory intended to appeal to an audience. Quantum theory's spooky action at a distance is not intended to comfort anyone. When Richard Feynman said that no one understands quantum mechanics, he identifies exactly the nature of science. It leads wherever the evidence and the theory takes us, even if it goes where we can't understand. That's why the imagination of science is so much wilder than the generally predictable, pedestrian, often imitative, imagination of the arts. Art is made for an audience. Science is free of that constraint. 

The proof of the pudding is that when a scientific program bends to an audience, it's no longer science. I hope that's not a circular definition, but if it is, then I think we all probably agree on that definition. It's not as if people like spinning things and that's why they came up with such a spinning earth theory. Same with the wave interference in the double slit experiment. The independence of the audience is a necessary condition of scientific investigation. 

Art is for audiences. Religion is for a community. Metaphors are made to help people understand a relation, and often to bias one aspect over another, the rhetorical means of motivating an audience. Metaphysics and historicism attempt to provide absolute truths to the reasoning mind. 

Science, however, is not addressed to any audience. Scientific investigation follows wherever the evidence or the conjecture leads. It may lead to an incomprehensible quantum world of spooky action at a distance so that one of its greatest exponents would say that nobody understands it. It leads to a round, spinning earth, not because that will entertain or attract the mind with mystery. It's not intended to be an awe-inspiring mystery. It's just a matter of empirical observation and theoretically justified conclusion. It doesn't generate mysteries in order to uplift or captivate anyone. Its purpose is to solve those mysteries. Dark matter is not a beautiful, awe-inspiring mystery of the ages. It's a problem to the current theory, a problem to solve, not to embrace with rapture. 

Of course, audience independence is not a sufficient condition. Churning out crazy theories of no interest to anyone does not qualify as scientific. The goal for the scientific investigation still has to be predictive.

So many different ways to view the world. How can we embrace them all if they conflict? Should we embrace them all? How can we compare or evaluate their merits? The common answer is, each perspective is useful in its own domain: religion for community and morality, spirituality for individual insight, metaphor for inspiration and explanation, imagination for innovation, art for all of these, and science for the material world and all the systems it includes, which includes all those other theories, spiritual, psychological, and metaphysical, and activities like art. But its intent is never to be captured by an audience. 

Whether it succeeds depends on the scientist's integrity. Every scientist knows that the investigation is supposed to be free of an audience. That's why funding is such a problematic for science, and sometimes a scandal.