Skip to content
← Back to blog
Engineering·September 29, 2026·7 min read

Claude Code runs with more autonomy than Claude's chat app, even on the identical model. Anthropic's own data says that's the interface talking.

Anthropic's June 2026 index: coding-agent sessions score 0.37 points more autonomous than chat, and still 0.26 points more with the model held constant.

Anthropic's own usage data answered a question most engineering leaders have been guessing at this year: when a company says it allows AI in the software development lifecycle, how much control is actually being handed over. The company's June 2026 Economic Index scored a sample of real Claude conversations for autonomy, on a scale from none to extreme, and found that sessions run through its Claude Code tool carry meaningfully more autonomy than sessions run through its own chat product, even when the underlying model is identical. If your team's AI policy names an approved model and stops there, this is the data saying that's the wrong thing to govern.

What Anthropic actually measured

Anthropic's Economic Index is the company's recurring look at what a sample of real Claude conversations gets used for, built by running an automated classifier over conversation transcripts rather than asking users to self-report. The June 26, 2026 edition, titled "Cadences," sampled conversations and artifacts created between April 10 and June 10, 2026, and Anthropic publishes the underlying dataset on Hugging Face under a CC-BY license, so the classification behind these numbers is checkable rather than taken on faith. One of the metrics the Index tracks, introduced in an earlier 2026 edition, is autonomy: how much of a task Claude completed with little ongoing human input, scored on a 1-to-5 scale from none to extreme. This is the first edition to break that score out by product surface, comparing sessions run through the Claude Code command-line tool against sessions run through Claude.ai chat and Cowork. It's a companion to our earlier look at the same report's March edition, which tracked coding work moving from chat into the API; this edition is the first to measure how much more autonomy that work carries once it gets there.

The gap that survives holding the model constant

The headline comparison: Claude Code sessions average 0.37 points higher autonomy than chat and Cowork sessions across the whole sample. For conversations that produced a script or code snippet specifically, the gap widens to 0.53 points. Part of that could be explained by Claude Code sessions simply running a more capable model more often, and the data shows they do: 54% of Claude Code sessions run on Opus, against 10% of chat and Cowork sessions. So Anthropic re-ran the comparison holding the model fixed, looking only at conversations that used Sonnet on both surfaces. The gap shrank, but it didn't close: Claude Code sessions still averaged 0.26 points more autonomy than chat sessions on the identical model. Anthropic's own reading of that result is direct: "the product used is likely more important than the underlying model."

What this costs a policy that only names a model

Most AI usage policies written this year answer one question: which models engineers are allowed to use. That's the easier half of the decision, because it maps cleanly to a vendor contract and a line item. Anthropic's data says the harder half, how much a given interface lets that model do without a human checking each step, doesn't move with the model at all. It moves with the surface. A team that approved "Claude Sonnet for engineering work" and considers the policy question closed has, by this data, actually made two different decisions depending on where an engineer opens that model: a comparatively supervised one in a chat window, where a person still reads and pastes each suggestion, and a substantially more autonomous one in an agentic coding tool that reads files, runs commands, and can stage a commit with far fewer human checkpoints per step. The model name on the approval line is identical in both cases. What it's permitted to do unsupervised is not.

Where the autonomy actually lives in your pipeline

This matters most at the exact point most engineering orgs haven't re-examined: the commit and merge gate. A policy that requires human review before an AI-suggested change ships assumes the AI's output arrives as a suggestion, the chat-shaped interaction the policy was probably written against. An agent that opens a terminal, edits several files, runs the test suite, and stages a commit has already completed several of those steps unsupervised before a human sees a diff at all. Anthropic's own figure for how much more autonomy that surface carries, even on the same model, is the first attempt anyone's published to put a number on how far that assumption has already drifted. The fix isn't a more conservative model. It's writing the autonomy policy against the surface an engineer is actually using, not the model card.

A worked example

Picture a 45-person, Series B developer-tools company that rolled out an agentic coding assistant to its platform team this spring, after the same underlying model had already been running for months on individual engineers' laptops as a chat-based pair programmer. Leadership's AI policy, written for the chat era, required a human-authored pull request and normal review for any AI-assisted change. Nobody rewrote the policy when the platform team switched to the agentic version, because it was, on paper, the same model behind a new interface. Two months in, an engineer let the agent handle a routine dependency bump overnight, unsupervised well past the point of opening a pull request; it also touched a config file outside the ticket's stated scope, in a way that passed CI but broke a downstream service the test suite didn't cover. The review gate everyone assumed was still in place had quietly moved from "before the change is written" to "after it's already running," and nobody had decided that on purpose.

Where this fits, and where it doesn't

A formal, surface-specific autonomy policy earns its cost once an agentic tool is touching shared infrastructure, holding repo write access, or running unsupervised long enough that a mistake compounds before anyone looks, exactly the profile Anthropic's data says an agentic coding session already carries more of than a chat session, model held constant. It's the wrong first investment for a two- or three-engineer founding team where every change, AI-assisted or not, already passes through one person's eyes before it ships; at that size the review is still happening, it just hasn't been written down as policy yet. The honest sequencing is the one that applies to most AI governance questions: name the surfaces your team actually uses, rank them by how much they run unsupervised, and gate the ones carrying the most autonomy first, rather than writing one policy that reads sensibly for chat and silently stops covering the tool that replaced it. Our look at how long an AI agent can be trusted to run before it needs a checkpoint and our piece on why agents fail once they reach production both feed the same design question this data raises: not whether to allow autonomy, but where to put the checkpoint once you already have.

If your team is past the point where one person eyeballs every change and needs that gate built properly rather than assumed, our Silicon Valley team scopes exactly this kind of production-engineering work under build your team; talk to us about which surface in your own pipeline is actually carrying the most unsupervised autonomy today.

Sources

Frequently asked questions.

Anthropic's Economic Index is a recurring analysis that classifies a sample of real Claude conversations against a standard task taxonomy, built from an automated classifier rather than user self-reports. The June 26, 2026 edition, titled 'Cadences,' sampled conversations and artifacts created between April 10 and June 10, 2026, and for the first time broke out its autonomy metric by product surface, comparing the Claude Code tool against Claude.ai chat and Cowork.

Anthropic's June 2026 Economic Index scored Claude Code sessions 0.37 points higher on average than chat and Cowork sessions on its 1-to-5 autonomy scale, widening to 0.53 points for conversations producing a script or code snippet. Anthropic attributes this to the interface itself: an agentic tool that reads files and runs commands completes more of a task unsupervised than a back-and-forth chat conversation does.

No, but it narrows. Anthropic's June 2026 report found that limiting the comparison to conversations using Sonnet on both Claude Code and chat still showed Claude Code sessions running 0.26 points more autonomously. Anthropic's own conclusion was that the product used is likely more important than the underlying model for how much autonomy a session carries.

Not by itself. Autonomy measures how much of a task ran without ongoing human input, not whether the outcome was wrong. The risk Anthropic's June 2026 data points to is a mismatch: a review policy written for a supervised, chat-shaped interaction being applied unchanged to a surface that is, by this same data, already completing more steps before a human sees the result.

Write the review and approval gate against the surface engineers actually use for a given task, not the model name on a vendor approval list. A team running an agentic coding tool with repo write access on shared infrastructure needs a checkpoint calibrated to how much of the work that tool completes unsupervised, which Anthropic's June 2026 data shows is measurably more than the same model handles in a chat window.