Security & Reverse Engineering(3)
The security focus is shifting from static LLM endpoints to complex, stateful agent systems, with Anthropic and Google DeepMind standardizing public discourse on agent jailbreaking.
Major AI labs, including Anthropic and Google DeepMind, are publicly releasing frameworks and disclosures on red-teaming autonomous agents.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.