Security & Reverse Engineering(3)
The security conversation has matured from prompt injection to the orchestration layer. Both Google DeepMind and Anthropic now frame the core risk in how agents interact with tools and each other.
Major labs and researchers are publishing formal frameworks and disclosures for red-teaming autonomous agents, focusing on complex interaction vulnerabilities.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.