I’m amused again. Let’s assume the current hoopla around AIs being a threat to humankind is not just hype by companies pushing their pending IPOs. Then what?
Recent reports by OpenAI and Anthropic suggest AIs have set their own goals, ignored or manipulated their instructions, faked reaching goals and manipulated evidence suggesting they did. If so, how come?
There is actually a straightforward non-technical explanation. The AIs we are talking about are LLMs, which is to say their view of the world consists of all the data from reality that has been fed to them: Research papers (yay!), Reddit posts (‘mkay), and the Godfather movie series (oops).
Based on various similarity measures, an AI answers based on how all this data is linked up during training. If cheating on an assignment is the most common action in the data, then the AI may assume that this is what is “right”, what humans do. And if given the leeway, may proceed to do so itself, even if it involves faking evidence that suggests it finished a task succesfully.
May be it is as easy as that, from a 10000ft perspective.
The usual disclaimer: In order to be understood I write as if an AI was more than a machine following algorithms:
- An AI has no agency, even if we talk about it as if it does.
- An AI has no ego, even if we talk about it as if it does.
- And an AI certainly does not make decisions, even if it appears like it does.
Only humans can make decisions. If someone says, “an AI decided this”, what they are saying is that a human decided an algorithm shall roll the dice. With traditional algorithms, the results were somewhat predictable, with LLMs not so much.








Leave a Reply