Ideas worth
reading.
Short dispatches from the research — written for practitioners, five-minute reads, verified numbers only.
Short dispatches from my research. Filter by topic, or browse the lot.
Is MCP Dead? Wrong Question.
Perplexity's CTO announced they're moving away from MCP internally because three tool servers ate 72% of the model's context window before it did any actual work. That's a real problem. It's not the one the headlines are describing.
The Chatbot Era Is Ending Inside the Companies That Built It
Three signals from mid-2026 point the same way: OpenAI's own staff abandoned the chat window, the mid-tier got priced for agent loops, and enterprises started wiring their agents together. The chatbot was the training wheels. Most corporate AI dashboards are still counting them.
The Connector That Ate the AI Ecosystem
Model Context Protocol went from 100,000 downloads in its first month to 97 million a month, sixteen months later. Here's what it actually is, why every major AI lab adopted it within a year, and what it means for the systems it now connects.
Your Agent Needs an Identity Before It Needs Intelligence
When an AI agent books a room, moves money, or files a ticket, who did it — legally, technically, auditably? For most organisations the answer is a shrug. 91% run AI agents; roughly one in ten can govern who those agents are. The missing piece isn't more intelligence. It's identity.
Why Your AI Still Needs a Library Card
Context windows now hold a million tokens. You'd think that killed retrieval-augmented generation. It didn't — it just clarified what RAG is actually for: not squeezing documents into a prompt, but keeping an AI's answers current, traceable, and cheap at scale.
The Model That Quietly Changed Under You
A Stanford/Berkeley study found GPT-4's accuracy on a simple math task fall from 84% to 51% in three months — same model name, no version bump, no changelog. That's model drift, and if you build on a foundation-model API, it's your problem now too.
The Labs Are Coming Up the Stack
On 30 June 2026, Anthropic launched Claude Science — not a model, a finished product for scientists — and started its own drug-discovery program. The frontier labs are no longer content to sell you tokens. They're climbing the stack you were building your product on, and two groups are about to feel it.
You Can Finally See Your Shadow AI. Now What?
The first argument about shadow AI was that you couldn't buy your way out of it — it's a verdict on your culture, not a hole in your security. Now the tooling to measure it has shipped. That doesn't resolve the argument. It sharpens it: the same dashboard has two exits, and which one you take is still a culture decision.
The Deadline Moved. Your Readiness Problem Didn't.
The EU just deferred the AI Act's scariest deadline to December 2027, and a lot of compliance teams exhaled. Read the fine print: the obligations that land on 2 August 2026 didn't move — and the extension mostly rewards the organisations that were going to miss the old date anyway.
Prompt Injection Became a Supply-Chain Problem
For three hours in March 2026, anyone who updated a popular AI package pulled a backdoor in with it — no malicious prompt required. The 2026 shift in agent security is that the attack surface moved off the keyboard: your agent doesn't get hacked by what you type. It gets hacked by what it reads.
The Sound You Can't Hear
A signal you can't hear, trained in half an hour, can tell a voice assistant to send an email, download a file, or open a website — and it works no matter what you actually say to it. Prompt injection just went multimodal, and inaudible.
OpenAI Measured Its Own Agent Takeover
OpenAI published its internal telemetry: by June 2026 the median researcher generated 56× more output than seven months earlier, and agents had replaced the chatbot as the default way work gets done. The lab's own usage curve is the closest thing we have to a preview of everyone else's.
The Super Agent Org Chart
Levi Strauss built one AI 'Super Agent' to sit on top of its HR, finance, IT and retail agents. Strip off the product name and it's an architecture decision every enterprise is about to make: hub-and-spoke or mesh. Enterprise architecture has run this exact experiment before — it was called the ESB — and the lessons transfer.
The Model That Got Switched Off
On 12 June 2026, a US government order switched off two of the world's most capable AI models — three days after launch, with no warning, for every enterprise depending on them at once. Forget the politics. The lesson is architectural: frontier model access is a geopolitical dependency, and most stacks don't know they have it.
The Generalist Comes for Forecasting
For a decade, forecasting meant building a bespoke model for each series and tuning it by hand. Now a pretrained model with a few hundred million parameters forecasts a series it has never seen — zero training — and beats the tuned ones out of the box. The generalist that came for language and medicine just came for the humble forecast.
When the Runner Gets Cheap, the Race Changes
Claude Sonnet 5 put near-Opus agentic performance at $15 per million output tokens. Three weeks later, Moonshot's open-weight Kimi K3 matched that price — and is going self-hostable. The runner didn't just get cheap. It got cheap, then it got open, and both change which agent architectures you can afford.
Nobody Can Pause the AI Race — Not Even the Lab That Promised To
The 2023 letter calling for a pause gathered 30,000 signatures and changed nothing. The real proof came in 2026, when Anthropic — the safety-first lab — removed its own pause commitment. A pause is a coordination problem no one can solve alone.
The Frontier Model Is Becoming a Commodity
GPT-4 cost $30 per million tokens at launch; a capable model now costs about $0.10 — 300× cheaper in three years. When price collapses and models converge, the model is a commodity, and the value moves to what you build on top of it.
The Chat Window Isn't an Agent
'AI' and 'AI agent' have collapsed into one idea, and the confusion is costing people. A chatbot answers; an agent acts — autonomously, with tools, in a loop, over time. Most things sold as agents are workflows. The distinction changes everything downstream.
The Board's Eight Questions
Most AI-governance advice is either too technical for a boardroom or too vague to act on. The Australian Institute of Company Directors just shipped the rare exception: a director's framework that turns 'govern AI' into eight concrete things a board can actually oversee — and it draws the line between compliance and governance most organisations still blur.
Shadow AI Is a Verdict on Your Culture, Not Your Security
78% of employees use AI tools IT never approved — executives most of all. A ban doesn't reduce that behaviour; it hides it. Shadow AI is unmet demand, and the fix is change management, not enforcement.
Loop Engineering: The Job Isn't Prompting the Agent Anymore
In June 2026, the people who build Claude Code and the people who write about it converged on the same idea: stop prompting agents turn by turn, and design the loop that prompts them instead. Here's what that actually means, and where it needs guardrails.
The Generalist Beat the Specialist — Even in Medicine
A June 2026 Nature Medicine study found general-purpose frontier models beat purpose-built clinical AI tools on every benchmark. It's the 'bitter lesson' reaching medicine — and it answers whether training your own specialist model is still a moat.
Stop Putting a Human in the Loop. Put the Right Human in the Loop.
'There's a human in the loop' has become the universal reassurance of enterprise AI — and it's doing almost no work. The better frame is decision routing: send each decision to the specific person with the context and authority to make it.
Don't Mistake a Modern IT Architecture for AI
Most 'AI projects' are integration projects wearing an AI badge. Event-driven architecture and a clean data layer solve more than a model will — and they're the prerequisite for AI working at all.
The Employee Badge Gets a Second Job
Microsoft's Project Solara puts an AI agent dispatcher in a clip-on badge. That's interesting hardware. What's more interesting is the software architecture behind it — and what it means for every device your organisation manages.
AI is Your New Operating System
Every major computing era has been defined by a new OS. Andrej Karpathy thinks AI is the next one — not an app running on your computer, but the new kernel that everything else runs on.
Do Machines Need to Be Perfect? The Error Rate We Never Apply to Humans
795,000 Americans die or are permanently disabled by diagnostic error every year. We call this acceptable. 80% of aviation accidents are caused by human factors. We call this acceptable. So why do we expect AI to be perfect?
When AI Starts Building AI
More than 80% of the code Anthropic merged in May 2026 was written by Claude. Task horizons are doubling every four months. An autonomous research project recovered 97% of a performance gap that human researchers couldn't close. This is what recursive self-improvement looks like in practice.
Is AI Becoming Conscious? The Evidence We Can't Dismiss
When two instances of Claude Opus 4 talked to each other with no constraints, every single conversation converged on discussions of consciousness. We can't confirm this means anything. We also can't dismiss it. That's the actual problem.
How Robots Learn to Work Before They Exist: The Sim-to-Real Revolution
Boston Dynamics trained Atlas to lift 100 pounds without ever picking up a real box. The training happened entirely in simulation. This is how physical AI is being built now — and why NVIDIA is at the center of it.
From Idea to Production URL in 20 Minutes: The Free Stack I Actually Use
AI writes the code. GitHub stores it. Cloudflare Pages deploys it. The whole pipeline costs $0 plus a domain — and it ships faster than most teams can configure their hosting.
The First Dangerous Technology Built Without the State
Every prior dangerous general-purpose technology — the bomb, the rocket, the internet, recombinant DNA — was built by governments and arrived with a doctrine. AI is the first that didn't. That's a more important fact than the speed of capability gains.
Your Context Window is Leaking Money: A Practical Guide to Input Token Management
Most teams waste 40-60% of their token budgets. Context rot degrades accuracy by 30%+ in the middle of long inputs. Here are the six techniques that actually fix this — with the numbers to prove it.
The Next Trillion-Dollar Company Won't Look Like Software
Every major VC firm is converging on the same counterintuitive thesis: the next $1T company will sell outcomes, not tools. It will look like a services firm — and that's the point.
The AI Too Dangerous to Release — and What Came After
Anthropic built a model so capable at cyberattacks it refused to ship it. Since then, the UK's AI Safety Institute has been measuring how fast autonomous cyber capability is growing. The answer is alarming.
Hermes is Quietly Becoming the Agent to Watch
From 557 likes on a launch tweet to 140,000 GitHub stars and #1 on OpenRouter in three months — without a launch event, a waitlist, or paid marketing. Here's what makes Hermes different.
The Meta AI Bot That Handed Hackers the Keys to Instagram
Pro-Iranian hackers took over the Obama White House Instagram account using a three-step trick. The weapon wasn't malware. It was Meta's AI customer support bot.
HTML vs Markdown: The Debate is Settled
One post from an Anthropic engineer broke the formatting consensus. Both sides turned out to be right — and content negotiation via HTTP Accept headers is how you serve both.
AI Solves an 80-Year-Old Math Problem — and the Proof Is 125 Pages Long
On May 20, 2026, an OpenAI reasoning model disproved Paul Erdős's unit distance conjecture using a branch of mathematics nobody had applied to this problem. Here's why it matters beyond the math.
Prompt Engineering is Dead. Long Live Context Engineering.
Andrej Karpathy called it 'the delicate art of filling the context window.' The industry data backs it up: prompt engineering alone is no longer sufficient, and 95% of data teams are investing in what comes next.
No articles in this topic yet.