Security & Reverse Engineering(3)
The field is rapidly moving from theoretical agent security to applied practice, with both Anthropic and Google DeepMind establishing public norms for vulnerability disclosure and testing.
Major labs are publicly disclosing agent jailbreaks and releasing red-teaming frameworks, while independent researchers are beginning to pentest agents in the wild.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.