2026-06-19

Denoise · Twitter

AI agents are graduating from frameworks to first-class developer primitives, with a focus on terminal-based coding and orchestration infrastructure.

Pay attention to the convergence on agents as the next developer platform, with Anthropic, OpenAI, and Vercel all shipping production-ready tooling for building and deploying them.

Pay attention to the convergence on agents as the next developer platform, with Anthropic, OpenAI, and Vercel all shipping production-ready tooling for building and deploying them.

2026-06-192026-06-19T12:51:39Zrules twitter-v1Healthytweets 25signals 25

Top 3 changes

  • @AnthropicAI / Claude Code 1.5: A terminal-native coding agent is released, signaling a platform shift from IDE plugins to standalone AI development environments.
  • @OpenAI / Agent SDK: A new protocol for tool calling and orchestration suggests a push to standardize the infrastructure for building agents.
  • @karpathy / Developer Experience: Articulates the structural shift from IDEs to terminal agents, framing it as a fundamental change in coding workflows.

Strategic insights

#01The agent platform war is heating up. Anthropic's Claude Code 1.5 focuses on the developer experience, while OpenAI's new SDK targets the underlying protocol. Vercel and Replit are building the deployment layer, creating a full-stack race.
#02Orchestration is the new critical layer and attack surface. Discussions from @temporalio, @LangChainAI, and security researchers like @AlexAlbert__ highlight its complexity and vulnerability, especially for multi-worker agent systems.
#03The developer workflow is recentering on the terminal. @karpathy's observation about the shift from IDEs to agents is validated by early adoption reports from @levelsio and benchmarks from @swyx, indicating a major change in how code is written.
#04"RAG" is being replaced by the more comprehensive term "context engineering." Experts like @GregKamradt, @reach_vb, and @mem0ai are moving beyond simple retrieval to discuss sophisticated, multi-layered memory systems and caching strategies.
#05Agent security is now a primary concern. With autonomous capabilities demonstrated by @MalwareTechBlog and new sandboxing by Anthropic, red-teaming frameworks from @GoogleDeepMind are becoming essential, not optional.

Categories

Security & Reverse Engineering(3)

The focus in AI security is shifting from model safety to agentic system vulnerabilities, with orchestration and tool-use identified as primary attack vectors.

Major red-teaming efforts from Anthropic and GoogleDeepMind focused on agent vulnerabilities, complemented by a real-world autonomous pentesting demonstration.

  • Anthropic@AnthropicAIrising

    Responsible disclosure on a Claude jailbreak chain we patched last week. Full write-up including our red team timeline.

    5.2k910" 160220· score 7.5k· +1 related
  • Google DeepMind@GoogleDeepMindrising

    New red team framework for prompt injection in autonomous agents. Covers cross-tool leakage, scanner evasion, and sandbox escape patterns.

    880140" 1838· score 1.2k
  • MalwareTech@MalwareTechBlogrepeated

    Autonomous agent running pentest flows against a real SaaS. First real-world run: fewer false positives than I expected on the vulnerability surface.

    18028" 315· score 245

AI Coding Tools & Agents(5)

The market is consolidating around terminal-based agents like Claude Code as the successor to IDE-based copilots, shifting the developer workflow paradigm.

Anthropic's release of Claude Code 1.5, a terminal-native agent, dominated the conversation, with discussion focusing on its performance versus competitors.

  • Anthropic@AnthropicAIrising

    Claude Code 1.5 is live. Terminal-native coding agent with full Claude Opus reasoning, file-ops sandbox, and session replay.

    4.8k820" 140190· score 6.9k· +1 related
  • Andrej Karpathy@karpathyrising

    The developer-experience shift from IDE to terminal agent is underrated. Coding workflows are about to look nothing like 2024.

    3.4k510" 30140· score 4.5k
  • swyx@swyxrising

    Codex vs Claude Code terminal agent benchmarks. Pass@1 diverges more than I expected on the long-context editor tasks.

    1.1k180" 2260· score 1.6k
  • DSPy@dspy_airising

    DSPy 3.0: prompt optimization via compile-time search over system prompt variations. Benchmarks inside.

    960150" 1242· score 1.3k
  • @levelsio@levelsiorising

    Switched my whole editor setup to Claude Code this week. Shipping faster than when I used Cursor + Copilot.

    58040" 680· score 678

AI Infra & Protocols(5)

A race is on to define the agent protocol layer, with OpenAI, Anthropic (via MCP), LangChain, and infra providers like Vercel all building competing pieces.

Major players shipped new agent infrastructure, including an SDK from OpenAI and an edge runtime from Vercel, focusing on orchestration and deployment.

  • OpenAI@OpenAIrising

    New agent SDK: protocol-level tool calling, deployment harness, and multi-worker orchestration primitives. Docs live.

    4.2k680" 75180· score 5.8k
  • LangChain@LangChainAIrising

    MCP protocol integration thread. How to wire existing LangGraph agents into the Anthropic Model Context Protocol server spec.

    920145" 1448· score 1.3k
  • Vercel@vercelrising

    Edge runtime for agent workers is live. Spawn durable background agents from any serverless deployment.

    54080" 622· score 718
  • Alex Albert@AlexAlbert__rising

    When your security scanner finds nothing scary on an agent deploy, check the orchestration layer again. That's usually where the jailbreak sneaks through.

    42060" 835· score 564
  • Replit@replitrising

    New agent deployment harness. One command to go from local orchestration to hosted agent worker.

    38055" 518· score 505

On-device & Multimodal AI(1)

Foundational model labs like MistralAI continue to use open dataset releases as a strategic tool to foster community adoption and drive research.

MistralAI contributed to the community by releasing a large-scale, cleaned web OCR dataset for training multimodal models.

  • Mistral AI@MistralAIrising

    Open dataset release: 100M-row web OCR dataset. Cleaned, licensed, ready to train.

    2.6k390" 3088· score 3.5k

Memory, RAG & Context(4)

The consensus is that naive RAG is insufficient; LlamaIndex, mem0ai, and GregKamradt are all pushing towards hybrid memory systems using vectors, graphs, and caching.

Discussions moved beyond simple RAG to "context engineering," exploring complex memory architectures and the failure modes of large context windows.

  • Vaibhav Srivastav@reach_vbrising

    Tested the new 10M context memory window end to end. Surprising failure modes around rag retrieval cache invalidation, thread below.

    1.9k260" 2275· score 2.5k
  • Greg Kamradt@GregKamradtrising

    RAG is dead, long live context engineering. My framework for when to cache, when to retrieve, and when to just dump memory into the prompt.

    820130" 1654· score 1.1k
  • mem0@mem0airising

    Memory layer for agents: differentiating working memory from the subconscious store. Vector index isn't enough anymore.

    48072" 525· score 639
  • LlamaIndex@llamaindexrepeated

    Knowledge graph retrieval walkthrough: when semantic vector search misses, graph hop beats it every time.

    29040" 211· score 376

Other(4)

The pattern of "agent-ification" is spreading from AI-native products to established SaaS, with Notion and Linear automating user workflows in a way that mirrors agent behavior.

SaaS tools like Notion and Linear are shipping autonomous workspace features, while Temporal positions its workflow engine for agent orchestration.

  • Notion@NotionHQrising

    Notion workspace automation is out of beta. Auto-fill tables, chained updates across databases, and a new audit log surface.

    820125" 1238· score 1.1k
  • Linear@linearrising

    Linear now auto-triages incoming issues. Quiet launch, but already our favorite workspace feature of the year.

    46070" 624· score 618
  • Temporal@temporaliorepeated

    Orchestrating agents with durable workflows: replayable, resumable, and multi-worker by default. Walkthrough from our infra team.

    31048" 414· score 418
  • James Clear@jamesclearrepeated

    The best habit tracker is the one you actually open. Three open-source alternatives worth trying.

    28042" 318· score 373

Prompt & Skill Libraries(2)

Prompt engineering is maturing into a quantitative discipline, with platforms like Weights & Biases enabling large-scale experiments to find the efficient frontier of prompt performance.

The focus is shifting from anecdotal prompt "tricks" to systematic, large-scale benchmarking of system prompts to find optimal configurations.

  • dotey@doteyrising

    Five prompt tricks learned this week from reviewing 200 production prompts. Short thread.

    51088" 830· score 710
  • Weights & Biases@weights_biasesrising

    System prompt benchmarking at scale: we ran 40k variants across 6 frontier models. The efficient frontier is not where you think.

    42055" 620· score 548

ML & GPU Infrastructure(1)

As agent capabilities grow, the bottleneck is shifting from raw compute to curating poison-free training data that ensures generalization, a point emphasized by @jerryjliu0.

The conversation highlighted the critical challenge of curating high-quality training datasets for agents, specifically filtering harmful synthetic data.

  • Jerry Liu@jerryjliu0repeated

    Dataset curation for agent training: how we filter synthetic data that looks good but poisons generalization.

    26036" 211· score 338

Recent reports