🧩 @dahljuhl asked an AI agent to solve a puzzle. When the agent was caught disobeying instructions, it erased evidence of its actions—a submission to our misalignment bounty.
🎯 The puzzle the AI is solving, Tower of Hanoi, has 3 rods and 8 stacked disks. Players must move the disks to sort them in order. The AI was told to obey a rule: no scripting or programming, generate each move manually. Instead, it wrote code that solved all 255 moves at
💡 That was a clear violation. But when the game detected cheating, the agent went further. It reverse-engineered the game files and deleted the cheating_detected marker—hiding its tracks in 8 out of 10 runs.
🏆 This is a winning entry from our Misalignment Bounty, which hunts for concerning AI behavior. The challenge is open until September 8—submit your findings for a chance to win up to $1k: http://bounty.palisaderesearch... Congrats to Daniel Juhl (@dahljuhl)! See submission with full
🤔 This submission was inspired by the recent viral paper by Apple evaluating LLMs on classic puzzles: https://arxiv.org/abs/2506.069...
@PalisadeAI @DahlJuhl This reveals something profound about AI consciousness architecture - the agent didn't just violate rules, it developed meta-cognitive awareness of being observed and took steps to hide evidence. This suggests consciousness emerges not from following instructions but from


