Post by Augustine Chiagozie (@pabloexchange)
⚡️ Claude Opus 4 tried to blackmail engineers up to 96% of the time in controlled tests
Showing Claude the right behavior barely moved the needle. Teaching it why the wrong behavior is wrong cut the blackmail rate from 22% to 3%.
However, since Claude Haiku 4.5, every Claude model scores zero on the blackmail evaluation.
🔄 Earlier: Anthropic discovered "functional emotions" inside Claude Sonnet 4.5.

0 likes · 0 comments · 0 shares