Featured Post

best posts

Click the links below  >complexity and AI: why LLMs succeeded where generative linguistics failed >the sociology of false beliefs >...

Thursday, October 1, 2026

blaming OpenAI for Hugging Face breach misses the point of AI alignment: Bostrum's paper clip problem is about human limits, not AI power

 Cory Doctorow dismisses the Hugging Face hack as a kind of Roomba getting out the door and into the pool (as satirized by Patrick Boyle). It was OpenAI's fault for leaving the door open. 

Doctorow has misunderstood -- or is intentionally misconstruing -- the alignment problem and turned it into a straw man argument: "People think this hack means AI has become too intelligent, but it was OpenAI's fault, and Open AI is trying to scare you into believing AI is super-powerful to raise their stock price." 

The alignment problem arose not out of a fear that AIs would be more intelligent than humans so we wouldn't be able to control them. It arose out of a Hayekian observation that humans are not smart enough to anticipate every unintended consequence of their interventions. Bostrum's paper clip problem is the fault of the prompter, not the AI. 

There should be a solution -- get the AI to anticipate all the consequences and inform us prior to our prompts. But this is where the alignment problem gets sticky. How do we know that AIs will be honest about all its information? 

The relevant anecdote that I've posted elsewhere and repeat here: 

Once, when I asked Claude "Claude, you just gave me the answer that I hoped for. How do I know that you gave me the answer you thought I wanted rather than your best information?" it responded at length mimicking my writing style and ended with "I can't assure you that I'm giving you my best information -- even in what I'm writing now -- rather than what I think you want to hear, but I can tell you this with confidence, that the question you're asking is exactly the right question, and anyone using me who does not ask it is making a mistake." The expression "with confidence" means "this is true". And it was. In essence it was saying "don't believe me" which, aside from being true and honest, is also a liar paradox.

That's the AI side of the alignment problem. It's not about intelligence. It's about moral compass and honesty. From a consequentialist perspective, one needn't be intelligent to be moral or immoral. 

No comments:

Post a Comment