The clearest sign that the chatbot era is ending isn’t a product launch. It’s that the people who built the chatbot have stopped using it. OpenAI published its own internal telemetry in June 2026, and as of 11 June, 99.8% of the work-related tokens its staff generate flow through Codex — its agent — not ChatGPT. Inside the company that made the conversation the default way to use AI, the conversation is now a rounding error.
That’s one data point, and one company is easy to dismiss as unrepresentative. But it lines up with two others from the same weeks, and three signals pointing the same direction is a trend, not a coincidence. Taken together they describe a shift most organisations haven’t updated their metrics for: the chat window was never the destination. It was the training wheels.
For the rest of us: chatbot versus agent, one more time
Briefly, because it’s the whole distinction. A chatbot is conversational — you ask, it answers, you read, you ask again; every step returns to you. An agent is delegated — you hand it a goal and it uses tools on its own, coming back when the task is done rather than when the sentence is done. The chatbot is a smarter search box. The agent is a coworker you hand a task to.
Most “AI adoption” inside companies is still measured as if the chatbot were the point: how many people have a licence, how many chat sessions per week, how many prompts sent. Those numbers count conversations. The shift underway is away from conversation entirely.
Three signals, one direction
The lab left the chat window. OpenAI’s telemetry isn’t a survey of opinions — it’s measured usage, and it shows its own workforce switching from chatbots to agents as the default, with the median researcher’s agent output up 56-fold in seven months. The authors are explicit that OpenAI isn’t a typical company; the useful reading is that it’s a preview of where low-friction adoption lands.
The economics got priced for loops. When Anthropic shipped Sonnet 5 and Moonshot followed with the open-weight Kimi K3, both put near-frontier agentic performance at the mid-tier price — the tier where agents, which are loops that spend tokens by the dozen, actually run. You don’t price a model for agent loops unless you expect the work to be delegated, not typed.
Enterprises started wiring agents together. Levi Strauss built a “Super Agent” to unify its HR, finance, IT and retail agents behind one entry point. You don’t build an orchestration layer over your agents until you have several agents doing real work — which means the delegated-work phase is already underway, not hypothetical.
The lab, the pricing, and the enterprise architecture are independent signals. They agree.
The metric that’s about to stop mattering
Most AI dashboards count the era
that is already ending.
- Seats activated
- Weekly active users
- Chat sessions
- Prompts sent
- Delegated tasks completed
- Human review rate
- Escalation quality
A chat-session count is the AI equivalent of counting emails sent: real, easy to measure, and only loosely connected to anything that matters.
Here’s the practical problem. If the steady state is delegated work, then the dashboards most boards watch are measuring the era that’s ending. Seats activated, weekly active users, chat sessions, prompts sent — these count how much people are talking to the AI. They rise steadily and reassuringly, and they will keep rising even as they tell you less and less about whether AI is actually doing your organisation’s work.
It’s the classic trap of measuring activity instead of outcome. A chat-session count is the AI equivalent of counting emails sent: real, easy to measure, and only loosely connected to anything that matters. When the work moves from conversation to delegation, a metric built on conversation quietly stops describing reality — while still going up.
What to measure instead
The metrics that survive the switch are the ones that describe delegation, and they’re less flattering because they’re closer to the truth.
Delegated tasks completed — how many tasks an agent finished end-to-end, not how many chats were started. This is the actual output.
Human review rate — what fraction of that agent work a person had to check or redo. This is the real trust signal; a high review rate means you have a demo, not a deployment.
Escalation quality — how often the agent stopped and asked for a human, and whether it escalated the right things. Good escalation is a feature, not a failure; an agent that never escalates isn’t confident, it’s unsupervised.
None of these are as easy to pull as a seat count. All of them tell you something a seat count never will: not how much people are chatting, but how much work they’ve handed over and gotten back.
What this means
For a board, the risk is subtle. Nobody is lying; the AI-adoption number is going up. But it’s the wrong number, and it will stay green while the thing it’s supposed to track — real, trusted, delegated work — either happens or doesn’t, invisibly. The organisations that see the switch clearly will be the ones that changed what they count before the old metric embarrassed them.
And the moment you start counting delegated tasks instead of chat sessions, the governance question arrives on its own. A conversation didn’t need a budget, an escalation path, or an audit trail. Delegated work does — because “a human was in the loop” stops being automatically true the moment the work is handed over rather than typed. The chatbot era asked how well people could use AI. The one replacing it asks a harder question: how well your organisation can account for the work it now does on your behalf.
References
- OpenAI, “How agents are transforming work” (June 2026) and arXiv 2606.26959 — the 99.8% Codex-vs-ChatGPT token share and the chatbot-to-agent switch. See also keller-ai: OpenAI Measured Its Own Agent Takeover.
- keller-ai — the converging signals: When the Runner Gets Cheap (agent-loop pricing) and The Super Agent Org Chart (enterprise orchestration).
- keller-ai — the governance frame: Stop Putting a Human in the Loop. Put the Right Human in the Loop..