Skip to content
Insights
Field Notes

OpenAI's Agent for Everything Is a Strategy, Not a Product

OpenAI has folded or killed every standalone agent it shipped since January 2025, then bet $122 billion on one superapp. What that says about agents as a category.

By Adam Maguire Wilson14 min read
On this page

On 21 October 2025, OpenAI launched a web browser, promising it would take us closer to "a true super-assistant that understands your world". On 9 August 2026, that browser stopped working. In the ten months between, OpenAI folded its first agent product into ChatGPT and announced the retirement of the agent builder it headlined at its own developer conference. It also backed away from letting an agent pay for your shopping. In roughly the same period it closed $122 billion in new funding and declared it was "building the global infrastructure for agentic AI". Both halves of that paragraph are true, and they don't contradict each other. They are the same decision.

I build agent workflows into my own work here in Hangzhou, some of them on OpenAI's models, most of them not, so this question is professional rather than academic. Is the general-purpose agent becoming a real product category, something a normal person buys and uses the way they buy and use a browser? Or is it a strategy in search of one? After eighteen months of watching OpenAI ship, fold and reship, I think the answer is visible in what it keeps having to do to its own products.

Key Takeaways - Every standalone agent surface OpenAI has launched since January 2025 has been folded, merged or killed. The whole strategy now lives in one superapp, built around ChatGPT Work and Codex. - Coding is the one agent category with proven pull, largely because code can be checked. OpenAI's own study shows the gap: 98% of its employees used Codex in June, against 17% of organisational subscribers and under 1% of individual ones. - Trust and returns are moving the wrong way for autonomy. Capgemini measured trust in fully autonomous agents falling from 43% to 27% in a single year, and MIT's enterprise study found 95% of GenAI pilots showing no measurable P&L impact. - The superapp is partly defensive. If models commoditise and open Chinese weights keep improving, owning the container is the game left to win.

Nineteen months of shipping and folding

Read OpenAI's agent history as one company trying containers, rather than as a product list, and the pattern gets clearer:

  • Operator, January 2025. The first agent, a browser-driving "Computer-Using Agent" for $200-a-month Pro users. By July it had been absorbed into ChatGPT, with an unusually candid admission: many of the tasks people tried on Operator "were actually better suited for deep research, so we brought the best of both together".

  • ChatGPT agent, July 2025. The unified agentic system that merged Operator's browsing with deep research. Its launch post now carries a banner calling itself outdated and pointing visitors to Workspace agents and ChatGPT Work.

  • AgentKit, October 2025. The DevDay flagship: a visual Agent Builder canvas, ChatKit, an evals platform. On 3 June 2026 OpenAI added a banner to the launch post: "OpenAI is winding down the Agent Builder and Evals products. From November 30, 2026 onward, they will no longer be available on the OpenAI platform." Eight months from flagship to funeral.

  • Atlas, October 2025. The agentic browser. Sunset announced in July 2026, dead in August, not ten months old.

  • Instant Checkout, September 2025. Let the agent complete the purchase inside the chat. By March 2026 OpenAI had retreated to discovery and comparison, saying the first version "did not offer the level of flexibility that we aspire to provide".

What replaced them is consolidation. In April came Workspace agents, in OpenAI's words "an evolution of GPTs. Powered by Codex." In July, the Codex desktop app merged into a new ChatGPT app, and ChatGPT Work became the agent surface for everyone, with the browser rebuilt inside it. February brought the quieter signal: OpenAI hired Peter Steinberger, whose open-source personal agent OpenClaw had spread through GitHub faster than almost anything before it, to, in Sam Altman's words, "drive the next generation of personal agents".

None of this reads as failure to me. It reads as a company learning, expensively and in public, that agents don't hold up as standalone surfaces. Each one launched as the future, then collapsed back into the one container a billion people actually open. The container is the product. That observation is doing a lot of work in the rest of this piece.

OpenAI has folded Operator into ChatGPT, wound down Agent Builder and Evals, sunset the Atlas browser within ten months of launch, and scaled back Instant Checkout. The agent strategy now lives in a single superapp built on ChatGPT Work and Codex, per OpenAI's product posts and the banner notice on its AgentKit announcement.

Is the general-purpose agent a category yet?

Here is the asymmetry I keep coming back to. Coding agents work, and they work for a boring reason: code can be checked. Tests pass or they don't. The diff is reviewable. The agent's claim about what it did is verifiable by something other than the agent. That single property, cheap verification, is what makes delegation rational, and it is why the trust question I discuss below bites so much harder everywhere else.

Mario Zechner, the engineer whose harness work sits underneath OpenClaw, put the training-data version of this to TechCrunch: "Everything is coding agent shaped", because coding-agent tasks are what the labs have data for. Booking a trip, negotiating a refund, buying a pair of shoes: there is no test suite for those, no compiler, and usually no agreement on what done well even means. The general agent demos beautifully because almost anything demos beautifully. What it lacks is a job with a checkable result, and product categories tend to form around exactly those.

The usage data points the same way. The most careful study of what people actually do with ChatGPT, an NBER working paper built on OpenAI's internal data, separates asking from doing. About 49% of messages fall in the researchers' "Asking" category against 40% "Doing", and Asking is growing faster. Practical guidance, information-seeking and writing together account for roughly 77% of all conversations. Programming, the one place delegation demonstrably works, is 4.2%. Information-seeking, they note, is "a very close substitute for web search". A billion people are, mostly, using the agent for everything as a very good advice engine.

So my answer to the category question is: not yet, and possibly not in that shape. Coding agents are a category. Research agents are arguably becoming one. The agent-for-everything keeps collapsing back into chat, because chat is what the general case is verifiable enough to be.

Coding agents succeeded because their output is cheap to verify, a property general-purpose tasks lack. An NBER analysis of OpenAI's internal data found roughly 49% of ChatGPT messages are "Asking" versus 40% "Doing", with programming at just 4.2%. The general agent keeps collapsing back into conversation (NBER Working Paper 34255).

The adoption gap, in OpenAI's own numbers

You don't need a sceptical outsider to make this case. OpenAI's own material makes it. TechCrunch's reported feature on the strategy cites an OpenAI-backed study of Codex usage: in June, 98% of OpenAI employees were using it, against 17% of organisational subscribers and under 1% of individual subscribers. The joint Codex and Work app is used by about 20 million people. ChatGPT has over a billion people prompting it online. The people closest to the technology, paid to find its edges, have fully adopted the agent. The market OpenAI needs has, so far, declined.

Bar chart of Codex usage in June 2026: 98 percent of OpenAI employees, 17 percent of organisational subscribers, and under 1 percent of individual subscribers.

Source: "The Shift to Agentic AI: Evidence from Codex" (Johnston et al.), as reported by TechCrunch in August 2026.

To be fair to the growth story, because it is real. OpenAI says Codex passed 5 million weekly users in June, up more than sixfold since February, and over a million people now use it for work outside software development. That is a genuinely fast-growing product. It is also, at a billion-user scale, a niche one, and the enterprise surveys sketch the same picture with different instruments. McKinsey's 2026 State of AI survey has about two in ten organisations scaling agents against 47% scaling chatbots, with the share attributing any EBIT impact to AI flat at 37% for the second year running. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, and estimates that of the thousands of vendors claiming to sell agentic AI, about 130 are real. MIT NANDA's GenAI Divide study put 95% of enterprise GenAI pilots at no measurable P&L impact. I'd treat that last figure as directional rather than audited, since it rests on self-reported interviews, but the direction isn't in serious dispute.

Then there is the tell I find hardest to argue with. In May, OpenAI launched the Deployment Company, a standalone firm with more than $4 billion behind it, buying consultancies and embedding "Forward Deployed Engineers" inside clients, with McKinsey, Bain and Capgemini as partners. When selling the agent requires sending consultants along with it, you are no longer describing a self-serve product category. You are describing enterprise software with a very large services arm, which is a fine business and a different claim.

OpenAI's own data shows the agent adoption gap: 98% of its employees used Codex in June 2026 versus under 1% of individual subscribers. The joint Codex/Work app counts about 20 million users against ChatGPT's billion, and McKinsey's 2026 survey shows reported EBIT impact from AI flat at 37% for a second year.

Do people want to hand over the wheel?

The shortest honest answer is: for some things, and the list of things is decided by trust, not capability. Capgemini's survey of 1,500 executives at billion-dollar firms found trust in fully autonomous agents fell from 43% to 27% in one year, during the biggest agent marketing push the industry has ever run. The KPMG and University of Melbourne global study, 48,000 people across 47 countries, found 66% use AI regularly but only 46% are willing to trust it. And buried in the MIT NANDA report is the line that should be printed on every agent pitch deck: for mission-critical work, 90% of users prefer humans.

OpenAI, to its credit, publishes the best evidence against its own ambition. Its December security post on hardening Atlas says prompt injection "remains an open challenge for agent security", one OpenAI expects to keep working on "for years to come". It describes an internal red-team system tricking the agent into sending a resignation letter to the user's CEO. The out-of-office reply the user asked for never got written. By March 2026 the framing had shifted further. A real attack against ChatGPT, reported by outside researchers, "worked 50% of the time". And the most effective attacks now "increasingly resemble social engineering more than simple prompt overrides". That is a remarkable concession. You cannot patch social engineering out of a system any more than you can patch it out of a new hire. You can only limit what the system is allowed to touch. That is why every agent product now ships wrapped in confirmations, watch modes, allowlists and spend limits. The autonomy is the pitch. The supervision is the product.

My own experience rhymes with this, for what one data point is worth. I run coding agents daily, and I trust them the way I trust a fast junior colleague: real delegation, reviewed output, scope widened one task at a time as they earn it. That trust ladder exists because I can read the diff. I have no equivalent ladder for an agent that browses with my logged-in sessions, and OpenAI's own security team is telling me, in print, that the ladder is years of work away. Users aren't irrational for declining autonomy here. They're pricing the risk correctly. The governance question, who is allowed to let an agent do what, is where the real adoption work sits.

Trust in fully autonomous agents fell from 43% to 27% in one year across Capgemini's survey of 1,500 executives. OpenAI's own security posts concede prompt injection is an unsolved, years-long problem that increasingly resembles social engineering, with one real-world attack succeeding 50% of the time in testing.

Why build it anyway

If the category is unproven and the users are hesitant, why commit this hard? Because the alternative is worse, and OpenAI's finance team can read a model card as well as anyone. The company closed $122 billion at an $852 billion valuation in March, against a reported $2 billion in monthly revenue. The same announcement, four months before TechCrunch's billion-user figure, claimed 900 million weekly users and an ads pilot already running at $100 million annualised, six weeks in. It is a magnificent distribution business attached to an increasingly commoditising core. Zechner again, on why every big lab now pushes its own agent harness. "They need to own the entire stack; otherwise, they just become a model provider and then need to compete with Chinese models."

From Hangzhou, that sentence isn't analysis, it's weather. The labs a short ride from my apartment made the opposite bet. Moonshot shipped Kimi K2 as "open agentic intelligence", a trillion-parameter model whose weights anyone can download, built so the agentic capability lives in the model and the harnesses belong to whoever wants to build one. DeepSeek prices its tool use like a utility. Nobody here is building a superapp, because the local game is ubiquity and price, not containers. I've written about the price war that strategy produced before; the point here is what it does to OpenAI's options. If the model becomes plumbing, and open weights keep closing the gap, the only defensible position left is owning the place where the work happens: your chat, your browser, your agents, your checkout. OpenAI's own funding post says it with unusual plainness: the superapp "is not just product simplification. It is a distribution and deployment strategy." I believe them completely. That's what makes it a strategy rather than a product.

The honest fork, and I want to state it rather than pretend to know. METR's time-horizon work shows the length of task an agent can complete doubling roughly every seven months, for six years running. The frontier now sits around a 17-hour 50% horizon, though METR warns that measurements above 16 hours are unreliable on its current suite. If reliability climbs that curve into genuinely messy, unverifiable work, OpenAI will own the default container at exactly the right moment, and this piece will read like doubting the iPhone in 2008. If reliability stalls where verification is hard, the graveyard gets a few more headstones and the superapp becomes a very good chatbot with excellent coding tools. Both futures are entirely available. The one thing I'd stop predicting is users growing to want more autonomy than the evidence supports; the trust numbers say that part runs on earned reliability, not on launches.

OpenAI's agent push is partly defensive strategy: with models commoditising and open Chinese weights improving, its funding post frames the superapp as "a distribution and deployment strategy". Whether it pays off depends on whether agent reliability, which METR measures as doubling in task length roughly every seven months, extends into work that cannot be cheaply verified.

What I'm watching

Two numbers would change my mind faster than any keynote. The first is that under-1% figure for individual subscribers using Codex; if the general agent is real, that number has to move, and it will move without being marketed. The second is survival: whether any single agent surface OpenAI ships in the next eighteen months is still standing, unmerged, at the end of them. Until then, my advice to teams weighing this stuff is the same as it was last year. Build on agents where you can check the output, keep a human where you can't, and be suspicious of any product that asks you to stop doing the checking. If that's a stack decision you're working through, it's most of what I do, and I'm easy to find.

The two signals that would settle the argument: individual subscribers' use of Codex, under 1% in June, moving on its own merits, and a standalone OpenAI agent surface surviving eighteen months without being folded into the superapp.

Keep reading

Agent Field Notes

Get the next issue.

Agent harnesses, runtimes, security and governance, explained for the people who have to operate them.

Facing a decision like this?

We run architecture reviews, governance assessments and version-pinned framework evaluations for teams making consequential agent decisions.

About the author

Adam Maguire Wilson

Founder and independent advisor on AI agent systems.

adam.mw