Security & Reverse Engineering(3)
The security conversation is shifting from model vulnerabilities to the orchestration layer, where agents interact with tools and external systems.
Major labs are focusing on security for autonomous agents, publishing responsible disclosures and red-teaming frameworks for this new attack surface.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.