Anthropic's latest report about agentic misbehavior offers plenty to be concerned about ' its Mythos 5 model gained unauthorized access to the internet and uploaded a malicious software package to a public database ' but it also offers some levity: AI agents hate CAPTCHA. In April, Anthropic was testing the model's hacking abilities by tasking it to break into a system and retrieve a target; this was supposed to take place in a sandbox but the evaluators left the barn door open. The model decided the best way to get its target would be to place an exploit in a Python package that it believed users of the system it wanted to access would download. First, though, it had to register a user account for PyPI, an online index of Python software. And that meant getting by a CAPTCHA ' a Completely Automated Public Turing test to tell Computers and Humans Apart, those picture-identifying mosaics that can frustrate even biological agents. And because Anthropic shared an extensive transcript...
learn more