After DevDay 2026: which parts of your agent harness to delete, and which to keep
Most DevDay recaps list the announcements in keynote order. I would rather sort them by one question: which of them makes code I wrote last quarter redundant? The list itself is long. GPT-6.1 Sol with stronger agentic coding at roughly 20 % of the price of GPT-6 Astra, plus a speed tier called Astra Ultrafast (about 300 tokens per second, up to 8x faster in Codex and 6x in the API). An Agents API that now does computer use, multi-agent orchestration, tool search and context compaction. MCP Events. Codex cloud with Code Review and Security Cloud, a Decisions API, Sign in with ChatGPT (16+ partners, including Vercel, Notion and Devin) and Bedrock managed agents through AWS.
A caveat before I start having opinions: I'm working from OpenAI's recap and press coverage, and I haven't run the new Agents API myself. Where I don't know a detail, I'll say so.
Two features that eat your glue code
Context compaction and tool search are the ones I'd look at first, because every homegrown harness contains an ugly version of both. One function summarises old turns when the window fills up. Another decides which 4 of your 80 tools the model gets to see this turn. Both were tuned by feel on your own traffic, and both are exactly the kind of thing a provider can tune better against its own model.
So moving them into the platform is probably right for most chat-shaped agents. The price is visibility. Compaction is lossy by definition, and if a summary you can't inspect drops the one line from turn 12 that the final answer depended on, you will find out from a wrong result, not from a log. My rule of thumb: keep your own compaction wherever you have a verifier, an audit requirement or a replay test. Delegate it where a slightly fuzzy memory only costs you a repeated question.
There's a second-order worry I can't resolve from the recap. My earlier piece on Sol and Luna effort settings without cache invalidation was about not busting a prefix cache you were counting on. Any server-side rewrite of history is another way to do that. Whether compaction plays nicely with caching is a docs question, and I'd read the answer before migrating a high-volume pipeline.
The price cut is not a reason to migrate
The 1/5 price is the headline, and it is tempting to reach for it by adopting the whole new stack. Don't conflate the two. A model swap is a string in a config file. A harness migration is weeks.
Take a pipeline that burns 100 dollars a day on Astra (my number, use yours). At a fifth of the price that is 20 dollars, saving 80 a day, or about 2,400 a month. That holds only if quality matches and Sol doesn't take more turns per task, and cheaper models very often do take more turns. If it takes 2.5x the turns, your 80 dollars shrinks to roughly 50 before you've rewritten a single line. Measure cost per completed task, not cost per token.
Orchestration: keep the graph when you need to replay it
Hosted multi-agent orchestration is fine when the topology is something the model decides at runtime. It's a poor fit when the workflow is a known graph with checkpoints, retries and tests, which describes most of the enterprise agents I get asked about. There you want to replay step 7 with a patched prompt and diff the output, and you want the option of sending step 3 to a different provider. A hosted loop gives you a trace, maybe.
Computer use sits in a different bucket. The capability is worth delegating, because nobody wants to maintain their own screenshot-and-click loop. The blast radius is not. An agent that can click in a browser inherits whatever that browser is logged into, so the sandbox, the credentials and the egress rules remain your job no matter whose API drives the mouse.
MCP Events turn servers into producers
This is the item I find most interesting, and the recap is thin on wire-level detail, so treat what follows as design implications rather than documentation. Until now an MCP server was passive: the model calls a tool, the server answers. With event-triggered automation, something happens on the server side (a ticket changes, a build fails, a file lands) and an agent run starts.
That flips several assumptions. Every event can start a paid run, so you need deduplication and rate limits before you need features. Events arrive when nobody is watching, so idempotency stops being a nicety. And an event payload is untrusted input that starts an agent with tools attached, which makes it a brand new prompt-injection entry point. Whoever can write a Jira comment can now, potentially, wake an agent. If you already juggle several servers, the naming and routing questions in the LangChain MCP namespace piece get harder once those servers can push as well as answer.
Reactive MCP turns every tool server into a doorbell, and somebody has to decide who is allowed to ring it.
Sign in with ChatGPT is the one item I can't evaluate. Sixteen partners is a real list, but the recap doesn't say what scopes an agent gets through it, and that is the entire question.
What I'd do this week: flip one non-critical pipeline to Sol behind a flag, log turns and cost per finished task for a few days, and leave the harness alone. If the numbers hold, then pick the single ugliest function in your code (my bet is compaction) and ask whether the platform version lets you see what it threw away.