AI & Work

The New AI Literacy Is Knowing What Not to Automate

The report arrives without the afternoon that once went into it. Before counting the hours saved, an organisation might ask what else used to happen in those hours.

Imagine a support manager watching an AI system turn a week of customer complaints into a report. The recurring problems are grouped, the tone is measured, the recommendations look sensible. A task that occupied Friday afternoon is finished before lunch.

The first response is relief. The second is a proposal: have it run every week.

Perhaps that is the right decision. But the afternoon contained more than report production. Someone read the awkward exchanges. They noticed where the approved answer failed to answer the question. They began to suspect that three apparently unrelated complaints described the same fault.

The report can be compared with last week’s report. The understanding acquired while making it is harder to place beside a benchmark.

The demonstration establishes that a report can be produced. The adoption decision also determines how people will come to understand the problems it describes. Sometimes both improve. Sometimes the immediate gain conceals a slower loss.

AI literacy needs to include the ability to tell those situations apart. Knowing how to delegate is becoming inseparable from knowing what delegation changes.

Capability is only the first test

The movement towards delegation is visible, although its extent needs careful qualification. Anthropic’s June 2026 Economic Index describes growing use of long-running agentic tasks across its products, requiring methods that go beyond analysing ordinary chat. This is evidence from one supplier’s users, interpreted by a company with a commercial interest in AI adoption. It is not a census of organisations. Anthropic’s report nevertheless captures a consequential change in what users hand over: sequences of work, with choices inside them.

Capability measures require similar care. METR evaluates the duration of software tasks agents can complete at specified success rates, using human completion time as the reference. A 50 per cent success horizon describes performance on a task suite; it does not establish that an agent can reliably run an equivalent stretch of someone’s working day. METR’s methodology makes the reliability threshold explicit. An adoption decision must do the same.

Three questions follow. Can the system produce an acceptable result? Does it improve the whole workflow once checking, corrections and monitoring are included? Is that improvement worth the changes to learning, relationships and responsibility?

These questions can have different answers. A complaints report might become cheaper while the team becomes less able to diagnose complaints. Equally, removing its repetitive preparation might give the manager time to investigate a problem previously ignored.

Nor does automating a task establish that a job can disappear. The ILO’s 2025 occupational exposure index examines potential task exposure, rather than observed job losses, and identifies job transformation as the likelier overall effect. The ILO research leaves room for the organisational choices that a capability score cannot settle.

The work inside the work

Some effort produces only the requested output. Other effort changes the person doing it.

Call the latter productive friction: effort that builds a capability or relationship the organisation will need again. Reading complaints can teach someone how customers describe a fault before the company has named it. Drafting an explanation can expose a gap in the writer’s understanding. Discussing a disputed case can reveal that colleagues have been applying different standards.

Wasteful friction adds effort without a defensible contribution to the result or to future capability. Copying the same figures between systems rarely becomes an apprenticeship merely because a junior employee does it.

The distinction requires evidence. Anyone defending a manual process should be able to name what it teaches, show where that learning matters and consider whether a better method could teach it. Time spent is a poor proxy for experience gained.

There is a related problem in what gets lost when an answer becomes detached from its source. An output can travel while the conditions that gave it value become less visible. In work, those conditions include the understanding left behind in the person who produced it.

The practical aim should be to preserve that understanding economically. Automate the formatting of the complaints report. Let AI suggest clusters. Retain a rotating sample of direct case reading, including cases that resist the categories. Then check whether people still notice new failure patterns.

This is a stronger defence of learning than keeping Friday afternoon intact.

Who learns to notice the mistake?

The difficulty becomes sharper when today’s supervisors learned through work tomorrow’s recruits may never perform.

An experienced manager may recognise a misleading summary because they have handled hundreds of difficult cases. Their successor could inherit the responsibility to approve summaries without inheriting an equivalent route to understanding them. That is the oversight paradox: delegating execution may remove some of the practice through which competent supervision develops.

It is a plausible organisational risk, not an established law of AI use.

A 2025 study by Erik Brynjolfsson, Danielle Li and Lindsey Raymond offers substantial counterevidence to simple deskilling claims. Following the staggered introduction of AI assistance among 5,172 customer-support agents, it found a 15 per cent average increase in issues resolved per hour. Less experienced and lower-skilled workers improved speed and quality; the researchers also found evidence consistent with learning. The published study concerns one deployment, but shows how AI can make useful expertise more accessible while people continue doing the work.

Other findings expose the difference between assisted performance and independent capability. In a randomised experiment involving nearly 1,000 pupils at a Turkish high school, access to a general-purpose GPT interface improved maths practice performance but lowered subsequent unassisted exam performance. A version designed with teacher-informed tutoring safeguards largely removed that harm, without demonstrating an exam improvement over the control group. Bastani and colleagues’ research measured short-term learning in school, not professional deskilling. Its relevance is the separation of a completed exercise from what the learner can subsequently do.

A smaller 2026 experiment brings that distinction into programming. Judy Hanwen Shen and Alex Tamkin randomly assigned 52 programmers learning an unfamiliar library to work with or without AI assistance. The assisted group performed worse on a subsequent knowledge quiz, with no statistically significant overall speed gain. Their preprint is a short study of skill acquisition, not proof that experienced developers are losing established abilities. Exploratory patterns suggested that how participants engaged with AI mattered too.

Taken together, these studies support a design question: does assistance help people engage with the difficult part, or let them bypass it?

An organisation can keep its experts on the payroll while quietly dismantling the work that made them experts.

Avoiding that outcome requires an explicit replacement for the apprenticeship being removed. Supervised casework, simulations and explaining decisions before seeing an AI recommendation are possible routes. Their value should be tested through unfamiliar cases and independent performance, rather than inferred from training attendance.

Approval is an activity, not a safeguard

A human approval step looks reassuring on a workflow diagram. Its practical value depends on what happens before the click.

Consider the manager presented with a polished complaints summary. If the underlying exchanges are inaccessible, the categories unexplained and the review squeezed between meetings, approval may register little more than plausibility. The system can claim that a human checked the work even though the human had no realistic opportunity to discover what was missing.

Human presence means someone participated. Human control means someone could understand enough to intervene, and had the time and authority to do so.

Research gives little support to treating the combination itself as a guarantee. Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone’s 2024 meta-analysis of 106 experiments found that human–AI combinations improved on humans alone on average, but fell below the better of human or AI performance alone. Results varied by task. The underlying studies ran only to June 2023, so this is no verdict on today’s agents. It shows why human–AI collaboration needs to be evaluated against both standalone alternatives.

Automation bias, the tendency to over-rely on automated recommendations, is one problem oversight must address. Article 14 of the EU AI Act explicitly includes awareness of that tendency, understanding system limitations and the ability to override or interrupt among its provisions for high-risk systems. These are requirements within a defined regulatory scope, rather than rules for every office tool, but they offer a useful account of what effective oversight entails.

For a complaints workflow, that could mean reviewing source cases before reading the synthesis, making disputed classifications visible and testing whether reviewers catch consequential omissions. Review time belongs in the business case. If nobody can afford to inspect the output properly, the approval step has not solved the problem.

Better evaluators change the boundary

Human review can be expensive, inconsistent and wrong. Automated checks may catch defects more reliably, while AI-assisted training may create practice opportunities that an ordinary workplace cannot provide. A second system may find omissions the first missed.

Those possibilities should change delegation decisions when they are demonstrated. They also need evaluation. Two systems agreeing may reflect a shared blind spot; a convincing explanation may still depend on a false premise. For the complaints report, the useful test is whether the checking process finds a deliberately omitted serious complaint, not whether it produces an articulate endorsement.

NIST’s generative AI risk profile recommends independent evaluations proportionate to risk, defined oversight responsibilities and incident-response arrangements. Its guidance treats governance as continuing work around a system, rather than a final inspection of its prose.

Better evaluation could justify removing routine human checks. It could also reveal that a task needs closer supervision than expected. An organisation should be able to move in either direction without treating the return of human involvement as an embarrassing retreat.

Choosing a delegation boundary

A useful framework separates five arrangements: Do, Assist, Review, Delegate and Automate. These describe permissions, not levels of organisational maturity. They are a decision aid, not a validated scoring system.

The critical distinction is between authorising an output, authorising a bounded assignment and granting continuing permission to act.

Arrangement Where execution and control sit Example within the complaints workflow
Do A person performs the core work, particularly where context or judgment is still developing. A new manager reads cases and constructs an initial account of the problem.
Assist AI supports defined parts; the person directs and performs the substantive work. AI retrieves related cases and suggests categories; the manager investigates and writes the interpretation.
Review AI produces the work; a qualified person evaluates it before it is used or acted upon. AI drafts the report; the manager checks claims against source cases before circulation.
Delegate AI completes a bounded assignment under explicit permissions and escalation rules. An agent investigates this week’s specified cases, produces a report for the authorised internal audience and escalates unresolved discrepancies.
Automate A standing trigger initiates recurring execution; human involvement is normally limited to exceptions and system oversight. A weekly process prepares and distributes a tightly defined operational report, with tested checks, monitoring and a named owner.

The framework needs two decisions alongside the chosen arrangement: what execution is permitted, and how the ability to evaluate it will be renewed.

The first establishes the boundary. Define what counts as an error, who could be affected and whether damage can be reversed. Identify the context needed to judge the result. Name the person responsible for failures and the conditions that require escalation. Permission to summarise complaints should not silently become permission to issue refunds or change customer records.

The second addresses the less visible dependency. If direct casework develops the expertise needed for review, decide who will continue doing enough of it, or what tested alternative will replace it. The same task may therefore be delegated by an experienced worker and assisted for a trainee. The trainee needs a route to greater responsibility, rather than permanent confinement to manual work.

A consequential error that reviewers cannot reliably detect is a reason to withhold autonomy even when most outputs look excellent. An easily checked, reversible operation with little learning or relationship value may warrant automation immediately. Averaging these factors into a single score would conceal precisely the exception that should govern the decision.

Trial the proposed arrangement against the current workflow. Count correction time and downstream mistakes as well as speed. Check whether the team retains the ability to diagnose unfamiliar cases. Agree in advance what evidence would trigger a narrower permission: a missed serious complaint, a change in policy, an unexplained shift in classifications.

No team needs to occupy one position throughout its work. It can automate report assembly, review the interpretation and retain human ownership of the response to a distressed customer.

When the person is part of the outcome

Some work deserves stronger human ownership because the relationship or responsibility is part of what is being delivered.

A customer seeking a routine status update may prefer an immediate automated answer. Someone challenging repeated mistreatment may need access to a person who can hear the history, reconsider the policy and accept responsibility for a remedy. Fluent sympathy alone does not settle that need.

Likewise, a manager can use AI to prepare for a difficult conversation. Delegating the encounter itself changes who has listened and who has made the commitment. The value of human involvement here is connected to authority, continuity and answerability. It does not depend on claiming that machines will never understand emotion.

These are choices about the service an organisation intends to offer. They should be informed by the people receiving it and the workers delivering it. A preference for human contact cannot be inferred from tradition, any more than a preference for automation can be inferred from a shorter handling time.

What the saved afternoon is for

Return to the report that arrived before lunch. The time saving may be real. So may the opportunity to use that afternoon better.

The more revealing question is whether anyone has decided what better means. Investigating recurring failures would be one answer. Giving inexperienced colleagues supported practice would be another. Increasing the volume of reports without improving anyone’s understanding would deserve a less enthusiastic calculation.

AI literacy should include recognising those differences: understanding capabilities, evaluating outcomes and knowing when to extend or withdraw permission. At an organisational level, it also means knowing which forms of expertise must be replenished and who remains answerable when execution moves elsewhere.

A mature organisation should be able to explain why a task is automated, why another requires review and why a third remains human-led. It should know what evidence would change each decision. The proportion automated tells us much less on its own.

When the next demonstration turns an afternoon into a minute, ask what the afternoon used to teach, and where that learning will happen now.


Sources and further reading

  1. Anthropic. “Anthropic Economic Index report: Cadences.” Anthropic, 26 June 2026.

  2. METR. “Task-Completion Time Horizons of Frontier AI Models.” METR, updated 8 May 2026.

  3. Paweł Gmyrek, Janine Berg, Karol Kamiński, Filip Konopczyński, Agnieszka Ładna, Balint Nafradi, Konrad Rosłaniec and Marek Troszyński. “Generative AI and Jobs: A Refined Global Index of Occupational Exposure.” International Labour Organization, Working Paper 140, 20 May 2025.

  4. Erik Brynjolfsson, Danielle Li and Lindsey Raymond. “Generative AI at Work.” The Quarterly Journal of Economics, published online 4 February 2025; May 2025 issue, 140(2), 889–942.

  5. Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences, 25 June 2025, 122(26), e2422633122.

  6. Judy Hanwen Shen and Alex Tamkin. “How AI Impacts Skill Formation.” arXiv, preprint submitted 28 January 2026; revised 1 February 2026.

  7. Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone. “When combinations of humans and AI are useful: A systematic review and meta-analysis.” Nature Human Behaviour, 28 October 2024, 8, 2293–2303.

  8. European Parliament and Council of the European Union. “Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Article 14: Human oversight.” Official Journal of the European Union, 12 July 2024.

  9. Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall and Kamie Roberts. “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile.” National Institute of Standards and Technology, NIST AI 600-1, 26 July 2024.