JournalAI & Society

Field guide / 7

Only 0.4% of occupations showed AI use across three quarters of their tasks

An open index separates observed chatbot use from agent capability and shows why occupational exposure is not the same as job replacement.

Sep 9, 20267By ISH Team
Only 0.4% of occupations showed AI use across three quarters of their tasks
Advertisement

Only 0.4% of occupations showed AI use across three quarters of their tasks

AI adoption reports tend to compress a labor market into one clean percentage. The new Open Economic Index asks a harder question: how far does observed AI use reach into the tasks that make up an occupation?

The University of Michigan-led project maps public chatbot conversations to the US Department of Labor's O*NET task taxonomy. Separately, it builds a tool-use benchmark around occupations that appear frequently in the data. Its most striking result is not that AI touches many kinds of work. Only 0.4% of occupations showed AI usage across at least 75% of their tasks.

A system helping with one writing task is not evidence that it can perform an occupation. The index makes that gap visible. Its public data and code also give researchers a way to inspect the assumptions behind the headline.

What the index actually measures

The adoption analysis starts with WildChat-4.8M, a public collection of user interactions with ChatGPT. After filtering, 1,674,316 unique English conversations remained. A classifier marked 789,768, or 47.2%, as occupationally relevant. Of those, 668,380 mapped to at least one O*NET task.

The mapping pipeline is layered. gpt-oss-120b summarizes each occupationally relevant conversation. Qwen3-Embedding-0.6B retrieves three candidate O*NET tasks. GPT-4o-mini then keeps candidates with core overlap. This is not a direct survey of workplaces, and the labels are model-produced rather than human-coded at scale.

ONET supplies the occupational structure. Its database describes work activities, skills, and other worker and job characteristics across the US economy. The project used ONET 30.1; the O*NET Resource Center currently distributes newer quarterly releases.

The resulting index shows concentration in arts, business and financial work, programming and mathematics, and subject sciences. Writing, programming, reading comprehension, critical thinking, and active listening are among the dominant skills. Legal services had the lowest adoption among the white-collar groups reported in the paper.

Breadth is not depth

An occupation can enter an exposure chart because AI appeared in one mapped task. That says little about how much of the occupation is affected. The researchers therefore calculate task coverage inside each occupation. Only 0.4% reached the threshold of observed AI use across at least three quarters of tasks, compared with a 4% estimate in a previous closed index.

This is not a market census. WildChat users are self-selected, location is incomplete, and public chat behavior does not represent every employer or worker. Translation between taxonomies can also change the result. The authors found less education exposure and more protective-services exposure than a previous index, partly because of mapping choices.

Task depth is a better lens than a broad exposure label. It interrupts the familiar slide from “someone used AI for a task” to “AI can do this job.” We made the same distinction when examining early-career hiring in AI-exposed occupations: an observed labor-market change needs a mechanism, a comparison group, and limits, not just an exposure score.

A separate test of agent capability

The paper does not assume observed use equals capability. It builds a second benchmark from nine frequently represented occupations with suitable tasks and tools. Nearly 5,000 scenarios cover 130 tasks, 100 servers, and more than 1,000 tools. Roughly 3,200 multi-turn trajectories were valid for analysis.

The researchers tested Kimi-k2.5 through an OpenAI Agents SDK harness with virtual tools. Kimi-k2.5 also generated scenarios and played the tool and user roles. An LLM judge scored tool-call accuracy, workflow completion, grounding, autonomy, and follow-up quality.

In single-turn tests, about 60% of required tool calls were correct. Mean workflow completion was 4.08 out of 5, while grounding averaged 3.7 out of 5. Multi-turn collaboration improved workflow completion but reduced tool-call accuracy and grounding.

Interaction helped the agent get closer to the user's goal, but it also created more opportunities to call the wrong tool or rely on unsupported details. This does not establish workforce replacement. The setup uses generated scenarios, virtual MCP servers, one tested model, and model-based judging. The paper also estimates that the available MCP servers cover at most 33% of the relevant software commodities in O*NET.

How to use the index responsibly

The project works best as a research instrument, not a leaderboard. Its public repository exposes the chat-index pipeline, analysis code, benchmark harness, MCP server definitions, and O*NET data. The combined OpenEconIndex dataset makes the output inspectable.

For teams evaluating AI at work, four practices follow from the methodology:

  1. Measure at the task level. Record which steps are assisted, which remain manual, and which require approval.
  2. Keep adoption and capability separate. Logs show what people try. Controlled tests show whether a system completes the work reliably.
  3. Test collaboration explicitly. A good single-turn score does not predict what happens after clarification, correction, and changing requirements.
  4. Audit grounding and tool use, not only completion. A plausible result produced through the wrong system or unsupported data is still a failure.

The same standard applies to internal assistants built on services such as ish.chat or an API gateway such as api.ish.chat. A useful deployment dashboard should show task coverage, failed calls, human interventions, and unsupported claims. A single “adoption” number hides the decisions that matter.

Read the two measurements together

The paper presents two linked but different measurements. Public conversations reveal where people are trying AI. Agent scenarios reveal how a model behaves when tools and follow-up turns are required. Neither alone measures productivity, job loss, or autonomous replacement.

The 0.4% figure is less a forecast than a warning about categories. AI use can be broad while remaining shallow inside most occupations. Any claim about work should say whether it measures curiosity, assistance, reliable completion, or substitution. The Open Economic Index gives researchers a transparent place to start.

#AI labor#Open Economic Index#O*NET#WildChat#MCP
Advertisement

Keep reading

Related stories

Browse the archive