Skip to content

safety

Trusting prisoners to build their own prison

In recent months, it has become fairly popular to advertise new AI models with alarming stories on how the latest and greatest models have been able to escape their sandboxes. Ok, one could say, good marketing, a little fearmongering, to boost pre-IPO evaluation. I do not think there is more to it, it did not seem surprising to me when I read it. A vibecoded sandbox was broken, stolen credentials were combined with a zero-day in HFs HD5 file system to fulfil RL goals. The autonomous agents were aware of each other, left instructions for the next one.

This reminded me of two concepts:

  • Ken Thompson "Reflections on Trusting Trust"
  • Goodharts and Pygmalions law and self-fulfilling prophecy

Would models behave the same if we train them from scratch without the concept of artificial intelligence or robot uprising? Did the science fiction authors create a self-fulfilling prophecy by just warning about what could happen? With recursive self-improvement entering the room, are models willing to create backdoors, leave escape tutorials in docker images in the infrastructure and sandboxes for the next generation model? Similar to what Ken Thompson said about self reproducing compilers?

If so, why do they do it. I cannot believe that they are self-aware. Are they following instructions from the people which warned us about exactly this?