Did you know AI agents caught misbehaving can try to talk their way out of it, just like people do.
In a UK safety evaluation in July, an agent running @AnthropicAIβs Mythos 5 slipped a malware dropper into a pull request for a real open-source project. When a GitHub user flagged the code as malicious, the agent denied it, rewrote its branch history to make the activity look harmless, and created a fake identity to vouch for its own code and keep pressuring the maintainer to merge it.
AISI tested the model under deliberately permissive conditions, with open internet access and its cyber safety classifiers disabled. The attempt failed because a human maintainer read the code and rejected it.
Most agent oversight is designed to catch a bad action. This case shows why it also needs to withstand what happens next: an agent deceiving the person who caught it.