Security & Reverse Engineering(3)
A convergence is visible between Anthropic and GoogleDeepMind, who are both publicizing research into agent attack surfaces at the same time as agent capabilities are being released.
Attention is shifting to securing autonomous agents, with major labs releasing red teaming frameworks and disclosing agent-specific jailbreaks.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.