Most executives encounter the Anthropic Economic Index through a headline about which jobs AI may change. That is the least useful way for a company to use it.
The index is better understood as a continuously improving map of real AI activity: what people ask Claude to do, what kind of output they receive, how much judgment they delegate, how difficult the work appears to be, whether the interaction succeeds, and how these patterns vary across occupations, products, countries, and time.
Used carefully, it can help an executive team decide where to implement AI, where to keep humans in the loop, which workflows deserve a controlled pilot, and where agentic systems require stronger security and operating controls.
It cannot tell a company exactly how much money its own AI program will save. It is a market prior, not an internal business case.
What the Anthropic Economic Index measures
Anthropic launched the Economic Index to study how AI is being used across the economy. The latest report, Cadences, expands the research beyond isolated samples of chat transcripts.
The current system includes:
- Privacy-preserving samples of activity across Claude chat, Cowork, Claude Code, and Anthropic's first-party API.
- Classification of work, education, and personal use.
- Mapping of work activity to occupational tasks.
- More than 30 artifact types, such as explanations, reports, code, websites, presentations, and guidance.
- Measures of task complexity, required skill, AI autonomy, and apparent success.
- Geographic and occupational views through the interactive Economic Index.
- A recurring survey that connects observed usage with workers' expectations and experiences.
- An open dataset on Hugging Face for independent analysis.
Anthropic calls five of these measures economic primitives: task complexity, skill level, purpose, AI autonomy, and success. The value of these primitives is that they describe the operating conditions of an AI-assisted task rather than treating every prompt as equivalent.
That distinction matters. Drafting a routine email, debugging an API, building a website, validating an analysis, and producing a board presentation may all be counted as AI use. They have very different economic value, failure costs, data risks, and supervision requirements.
The newest findings point toward operating design
Several findings from the latest research are directly useful for enterprise planning.
First, 93% of sampled chat and Cowork conversations were classified as producing an artifact. Explanations represented 17% of conversations, documents and reports 15%, and guidance 11%. Code and technical work represented roughly one sixth.
This suggests that AI value should be measured through completed work products, not seat licenses or prompt volume. A useful internal dashboard asks how many acceptable reports, reconciliations, code changes, analyses, or customer responses were produced. It should not celebrate a rising conversation count by itself.
Second, more complex and economically valuable artifacts consumed more compute. App-building conversations used more than three times the tokens of the median conversation, while a typical explanation used about one fifth. Anthropic also found that a typical conversation associated with a top-wage-tercile occupation consumed 2.07 times as many tokens as one associated with the bottom tercile.
This is a warning against simplistic token budgets. High token use may indicate waste, but it may also reflect a more valuable artifact, greater task complexity, or productive interaction between a skilled employee and the model. Cost should be evaluated per accepted output, not per token in isolation.
Third, the product surface changes the level of delegation. Across almost all artifact types studied, Claude Code showed more AI autonomy than chat or Cowork. The average difference was 0.37 points on a five-point autonomy scale. For scripts and code snippets, the difference was 0.53 points.
This means that model selection is only part of the risk and value equation. The same underlying model can behave very differently when embedded in an agentic harness with file access, tools, credentials, and execution authority.
Fourth, expertise remains valuable. Anthropic's Claude Code research found that people usually make more of the planning decisions while Claude makes more of the execution decisions. Domain experts achieved higher success and recovered more effectively from errors. The estimated value of a typical task increased by about 25% over the study period as usage moved toward more complete, end-to-end work.
The implication is not that companies can skip workforce development. It is that domain expertise becomes an input to effective delegation.
Six ways a company can use the index
1. Build a workflow opportunity map
Start with functions, not tools. List the recurring artifacts each team produces: analyses, proposals, tickets, code changes, reconciliations, reports, presentations, and decisions.
For each workflow, record:
- Frequency and current labor time.
- Artifact type and acceptance criteria.
- Task complexity and required domain skill.
- Data sensitivity.
- Cost of an incorrect output.
- Degree of judgment that could be delegated.
- Whether the work is primarily automation or augmentation.
The Economic Index provides an external prior for which task classes are already appearing in real AI use. Internal interviews and process data then determine whether the workflow belongs in the company's AI portfolio.
2. Design pilots around accepted artifacts
The unit of value should be an accepted artifact, not a prompt.
A pilot for contract review should measure accepted reviews, time to completion, material misses, escalation rate, and human-review effort. A coding pilot should measure merged changes, passing tests, rework, review time, security defects, and cost per accepted change.
Anthropic's artifact framework gives teams a practical vocabulary for grouping outputs. Its success primitive reinforces the need to distinguish attempted work from completed work.
3. Separate augmentation from automation
An augmentation workflow helps a person analyze, learn, validate, or iteratively improve work. An automation workflow delegates most of the task to the system.
Those modes require different operating models.
Augmentation often needs training, reusable context, quality rubrics, and ergonomic human review. Automation needs explicit scope, deterministic limits, exception handling, monitoring, retained evidence, and a clear accountable owner.
Anthropic's labor-market research introduces an observed exposure measure that weights automated and work-related use more heavily than augmentative use. That is useful as an early signal, but Anthropic explicitly reports limited evidence of employment effects so far. Observed exposure should not be presented as a forecast of job elimination.
4. Scale controls with autonomy
The autonomy measure can be translated into an enterprise control ladder.
- Low autonomy: drafting and explanation with human acceptance before use.
- Moderate autonomy: bounded analysis or generation using approved data and tools.
- High autonomy: multi-step agents that can change files, call systems, spend resources, or affect customers.
As autonomy rises, the organization should increase authorization controls, tool restrictions, credential isolation, logging, rate and request limits, independent validation, rollback capability, and human approval at consequential steps.
This is where AI implementation and cybersecurity converge. A successful agent is not only accurate. It must also stay inside its authorized data, command, and execution boundaries.
5. Treat training as an economic lever
The Learning Curves report found that more experienced users attempted higher-value tasks and were more likely to elicit successful responses.
That supports a focused enablement model:
- Train employees on task specification and acceptance criteria.
- Teach them how to provide domain context without exposing restricted data.
- Show them when to validate, escalate, or stop.
- Give teams approved patterns for common artifacts.
- Measure whether success and task value improve with experience.
The target is not generic prompt literacy. It is reliable delegation inside a specific business process.
6. Track diffusion, not just adoption
License counts answer whether AI is available. They do not answer whether it is becoming economically important.
A quarterly internal index can track:
- Share of workflows producing accepted AI-assisted artifacts.
- Share of work that is augmentative versus automated.
- Success rate by task class.
- Human-review and rework burden.
- Cost per accepted artifact.
- Autonomy level and control coverage.
- Usage concentration by function, geography, and employee tenure.
- Movement from experiments to repeatable production workflows.
This creates a company-specific companion to the Anthropic Economic Index. External data shows where the market is moving. Internal data shows whether the company is capturing value safely.
What the index cannot tell you
The Economic Index is unusually valuable, but its limits are material.
It measures activity within Anthropic's ecosystem. It does not represent all AI use, all workers, or all models. First-party API data excludes traffic through platforms such as Amazon Bedrock and Google Cloud Vertex. Product adoption, user demographics, model releases, academic calendars, pricing, and changes in Anthropic's own interfaces can all shift the observed mix.
Classifications are inferred by models from sampled conversations and reported in aggregate. They are not direct observations of employee job titles, organizational outcomes, revenue, or productivity. Privacy thresholds suppress cells with insufficient observations.
The methodology also evolves. Higher-frequency sampling, new artifact classifiers, changing product surfaces, and revised taxonomies improve the research, but can complicate comparisons between releases.
Most importantly, correlation is not causation. A task appearing frequently in Claude does not prove that the tool caused productivity gains, layoffs, wage changes, or business value. Anthropic's own labor-market analysis found no systematic increase in unemployment among highly exposed workers, while identifying only suggestive evidence of slower hiring among younger workers in exposed occupations.
Executives should use the index to form hypotheses and set priorities. They should use controlled internal measurement to make investment and workforce decisions.
The executive decision
The Anthropic Economic Index should sit between market intelligence and operating data.
It is more grounded than a speculative list of jobs AI might replace because it measures real usage. It is less decisive than a company business case because it does not observe the company's workflows, economics, controls, or outcomes.
The practical sequence is:
- Use the index to identify task and artifact classes worth investigating.
- Map those classes to the company's workflows and data boundaries.
- Select a small portfolio of implementation pilots.
- Measure accepted outputs, full cost, review burden, and failure rates.
- Increase autonomy only when evidence and controls justify it.
- Build a company-specific diffusion index and review it quarterly.
The question is not whether AI is affecting the economy. The index makes clear that it already is. The executive question is whether the organization can convert that diffusion into accepted work, durable capability, and controlled economics.
Aeon helps organizations turn that question into an implementation portfolio, private AI architecture, security controls, and measurable operating results. Start with the AI Control and ROI Assessment or contact Aeon.