Security & Reverse Engineering(3)
Major labs like AnthropicAI and GoogleDeepMind are establishing security best practices for agents in parallel with their release, indicating the maturity of the ecosystem.
Red teaming frameworks and responsible disclosures are being released for autonomous agents, focusing on prompt injection, tool leakage, and orchestration-layer vulnerabilities.
Anthropic@AnthropicAIrising Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.
Google DeepMind@GoogleDeepMindrising New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.
MalwareTech@MalwareTechBlogrepeated Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.