The Thesis

Organizations are answering AI risk with human review, and human review is exactly the skill that AI use wears down first. Across medicine, software, and knowledge work, reliance on AI is measurably degrading the unassisted abilities that reviewing depends on, while confidence in one’s own judgment inflates in the opposite direction. Three forces compound.

Deskilling. Reliance on AI atrophies the underlying expertise a reviewer needs. The capability being lost is not the ability to produce the work; it is the ability to catch what is wrong with it.

Calibration collapse. People grow more confident in their judgment even as it gets worse, so the usual link between competence and self-doubt breaks. The signal that would normally tell someone to look harder stops firing.

Oversight mechanics. The machinery of review fails quietly at scale, from rubber-stamping under volume to the way AI explanations lull rather than sharpen scrutiny.

When a control depends on a capability that the controlled activity destroys, the control is weaker than the org chart suggests.

The Signal

Three developments worth watching this week.

Signal 01
1. The deskilling evidence is no longer anecdotal.

What happened. Two domains supplied measurement where there had been argument. In The Lancet Gastroenterology & Hepatology, Budzyń and colleagues tracked 19 experienced endoscopists across four Polish centres participating in the ACCEPT programme. After the clinics adopted AI polyp-detection tools, the doctors’ adenoma detection rate — the share of colonoscopies in which they find at least one precancerous growth — fell from 28.4% before AI to 22.4% when working without it. Separately, a randomized trial from Anthropic, posted January 29, 2026, gave 52 mostly junior software engineers a coding task; the group using an AI assistant scored 50% on a comprehension quiz taken minutes later versus 67% for those who worked by hand, a gap of roughly two letter grades, with the steepest deficit on debugging.

Why it matters. Debugging and unaided detection are the abilities a reviewer uses to catch an AI’s mistakes. Each study carries a caveat: the Lancet work is observational rather than randomized, and the Anthropic trial comes from a vendor studying the assistant it sells. But read alongside the 2026 Future Ready Healthcare Survey (Wolters Kluwer Health with Ipsos, June 2), in which 74% of clinicians reported concern about deskilling and growing overreliance on tools that reduce their ability to independently identify inaccuracies, the direction is consistent across independent sources and independent methods.

Second-order effect. The pipeline that produces future reviewers is thinning. Ford deployed roughly 900 AI-powered inspection cameras, then brought back 350 veteran “gray beard” engineers after the automated systems underperformed — a reversal that helped the automaker top the JD Power 2026 US Initial Quality Study for the first time in sixteen years. Its VP of vehicle hardware engineering, Charles Poon, conceded the company had “mistakenly” assumed AI plus existing requirements would yield a high-quality product. IBM moved the other way, announcing in February 2026 it would triple US entry-level hiring, with junior developers spending less time writing basic code and more time supervising AI output.

Signal 02
2. Confidence rises as competence falls.

What happened. A study by Fernandes, Villa, Welsch and colleagues at Aalto University, published in Computers in Human Behavior in February 2026, found that under AI use the Dunning-Kruger effect — the familiar pattern in which low performers most overestimate themselves — does not merely persist. It vanishes. Higher AI literacy correlated with greater overconfidence and lower metacognitive accuracy. Wharton researchers Steven Shaw and Gideon Nave, in a January 2026 paper introducing the term “cognitive surrender,” found across 1,372 participants that people accepted incorrect AI answers about 80% of the time and correct ones 93%, and rated their own confidence 11.7% higher than non-AI users, with confidence climbing even after errors. METR’s randomized trial found experienced open-source developers were 19% slower with AI tools while estimating afterwards that AI had made them 20% faster, having forecast a 24% speedup beforehand: a 39-percentage-point gap between perception and measured reality.

Why it matters. Calibration — the match between how well you think you are doing and how well you actually are — is the reviewer’s early-warning system. When it collapses, people stop looking precisely when they should look harder. Christopher Koch’s March 2026 preprint calls this “metacognitive decoupling,” a widening gap between output quality and the ability to judge one’s own understanding. The most AI-fluent staff, the ones organizations trust to supervise AI, are the ones this literature flags as most likely to be overconfident.

Second-order effect. Self-reported productivity and confidence become unreliable inputs for governance. The METR gap means a leader who asks reviewers whether they are keeping up will get reassuring answers the data contradicts, so any oversight metric built on self-assessment misleads by construction.

Signal 03
3. Oversight fails at the moment governance leans on it.

What happened. Harvard Business School’s Alex Chan ran an experiment with 2,512 participants acting as loan officers on real $10,000 decisions. Fewer than half chose to view the AI’s explanation; avoidance was worst when bonuses depended on the outcome; and when explanations were viewed they increased overrides, as summarized in HBR in June 2026. The finding cuts against the common assumption that explainable AI produces sharper scrutiny. Meanwhile the arithmetic of agentic deployment is unforgiving: one analysis estimated that 50 agents making 20 tool calls an hour generate 1,000 approval-eligible events hourly, so routing even 10% to humans means 100 approvals an hour, or over three full-time staff doing little but rubber-stamping. Gartner projects 40% of enterprise applications will embed task-specific AI agents by the end of 2026, up from under 5%.

Why it matters. Regulation is tightening its grip on human oversight just as the grip weakens. EU AI Act Article 14 requires that overseers be able to stay aware of the tendency toward automation bias and to override or disregard the output. The European Commission’s draft high-risk classification guidelines, published May 19, 2026, make clear that human involvement does not, in itself, exclude a system from the high-risk category, so a rubber-stamp reviewer does not lower the regulatory bar. The FDA’s revised Clinical Decision Support guidance, issued January 6, 2026, keeps its automation-bias warning and holds that software for time-critical decisions generally fails the test that lets a clinician independently review the basis for a recommendation.

Second-order effect. “Human in the loop” becomes compliance theater. When approvals arrive faster than anyone can weigh them, the checkbox is ticked and no one is actually watching, which converts a safeguard into a liability that surfaces only after an incident.

The Metric

The Skill Stops Being Practiced Adenoma detection rate, 19 endoscopists working without AI BEFORE AI 28.4% AFTER AI 22.4% CHANGE −6.0 points Same doctors. Same procedure. No AI in the room either time.
Both figures are unassisted colonoscopies. 28.4% is the detection rate in the three months before the four centres adopted AI polyp detection; 22.4% is the rate in procedures performed without AI after adoption. The study is observational, not randomized, so exposure to AI is associated with the decline rather than proven to have caused it. Source: Budzyń et al., The Lancet Gastroenterology & Hepatology, vol. 10, 896–903 (2025).

What it measures. The fall in experienced endoscopists’ unassisted adenoma detection rate after their clinics adopted AI, from 28.4% to 22.4%.

Why it matters now. It is the cleanest real-world number showing that AI exposure can degrade an expert’s unaided performance on the same task, with a direct line to patient outcomes. That degradation is precisely the mechanism every “a human will catch it” control assumes away.

Readers of Searchlight 017 will recognise this six-point fall, which appeared there in The Lens as one line from the 2026 International AI Safety Report. We return to it deliberately: what was a secondary signal eight days ago is the central evidence here, read from the primary study rather than the summary.

The Counter-Case

A brief that asks readers to test claims should mark the limits of its own. The evidence assembled here is convergent, not conclusive. The Lancet finding rests on 19 endoscopists in one country under an observational design, which establishes association rather than cause. The Anthropic trial ran 52 participants, came from a vendor studying the assistant it sells, and measured comprehension minutes after the task rather than durable skill. The Aalto and Wharton results are laboratory studies of self-assessment. METR, strictly read, measures a perception gap rather than skill loss at all, and belongs in this argument for what it says about the reliability of self-report, not as deskilling evidence in its own right. Each of the three forces is separately evidenced; their compounding is inferred rather than measured. What makes the pattern worth acting on is that independent methods in unrelated domains point the same direction, not that any single study settles it. Larger randomized work that fails to reproduce the effect should move this thesis, and we would report that.

The stronger objection is historical. Every automation has shifted the skill mix, and the shift is usually fine. Navigators lost celestial navigation, accountants lost mental arithmetic, and almost no one argues we should have preserved either. Skills atrophy because they stop being load-bearing, and calling that decay rather than specialization is often just nostalgia. The aviation precedent is instructive precisely because it drew the line rather than defending every skill: pilots stopped hand-calculating fuel burn without objection, and the FAA intervened only when manual handling decayed, because manual handling is what the emergency procedure assumes. That is the test worth importing. The question is not whether an AI-exposed skill is fading, but whether a control you already claim to have depends on it. Where it does not, deskilling is specialization and should be allowed to proceed. Where it does, and human review is currently the most widely claimed control in enterprise AI governance, the decay is an unbooked liability sitting underneath a safeguard that has already been promised to regulators, boards, and customers.

The Playbook

Six moves to keep the oversight capability your controls assume you have.

Step 01
Build a skills inventory mapped to AI exposure.

List the judgment-critical tasks in each role, mark which ones AI now performs, and treat any task that is both high-stakes and highly automated as a deskilling watch item. You cannot preserve a capability you have not named.

Step 02
Run AI-free assessment windows.

Periodically test whether staff can perform and verify core work unaided, because unaided competence is what your review control actually depends on and self-reported confidence will not reveal its decay. Gartner’s Daryl Plummer forecast in an October 2025 briefing that critical-thinking atrophy from generative AI use would push half of global organizations to require “AI-free” skills assessments through 2026.

Step 03
Engineer deliberate friction into review.

Give reviewers predefined acceptance criteria rather than a bare Approve button, require them to record a reason before they can align with the AI, and suppress the AI’s recommendation until the human forms an independent view. Chan’s evidence shows that making an explanation available is not enough when incentives point toward willful blindness.

Step 04
Prescribe a maintenance dose of unassisted practice, modeled on aviation.

After Air France 447 and a documented rise in manual-handling errors, the FAA issued SAFO 13002 (2013) and SAFO 17007 (2017), urging airlines to keep pilots hand-flying so automation-era skills do not decay. Health systems and engineering teams can borrow the template directly, scheduling regular AI-off practice for skills that must survive an AI failure.

Step 05
Budget verification workload before scaling agents.

Treat human review as a capacity with a ceiling. Model the approval events your agents will generate, route by risk and confidence rather than sending everything to one queue, and staff the review function honestly. One hundred approvals an hour is not oversight. It is fatigue.

Step 06
Redesign junior roles to protect the expertise pipeline.

Ford’s rehiring of “gray beard” engineers and IBM’s tripling of entry-level hiring point the same way: someone has to become the senior reviewer of 2035. Give juniors deliberate unaided reps and mentorship, not just prompt-and-ship work.

The Verification Test

Claim Under Test

“Our AI deployments are safe because a human reviews every output.”

Test. Pull a random sample of recent approvals and ask each reviewer to explain the basis for the decision without reopening the AI’s output. Measure how many can, how long review actually takes, and the approval rate per reviewer per hour.

Pass criteria. Reviewers can independently reconstruct the reasoning, override rates are nonzero and track genuine error, and review throughput is low enough to allow deliberation.

Fail smell. Near-100% approval rates, reviewers who can only restate the AI’s own explanation, batched sign-offs, and any queue where approvals per hour exceed what a person could plausibly consider.

The Lens — Horizon Search Institute

Human Performance

Skill and confidence are decoupling under AI use, which means performance management has to measure unaided competence and calibration directly rather than trusting self-report. The METR gap is the clearest warning: the people best placed to notice the decline are the least likely to report it.

Responsible AI

Explainability was meant to make oversight real. Chan’s loan-officer experiment suggests it does something narrower: explanations are avoided when incentives favor agreement, and change behavior only once someone has already chosen to look. Availability is not a control.

Governance & Diplomacy

The EU and the FDA are converging on the same position from different directions: a human in the loop does not lower the bar unless that human can demonstrably act on what they see. Article 6 classification and the revised CDS guidance both turn oversight from a box into an evidentiary claim.

Links Worth Your Time

Sources
  1. Lenharo, M. Is AI ruining our skills? Early results are in — and they’re not good. Nature, June 18, 2026 (corrected June 22); reprinted Scientific American, July 5, 2026.
  2. Budzyń, K., Romańczyk, M., et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology, vol. 10, 896–903 (2025).
  3. Shen, J.H. & Tamkin, A. (Anthropic). How AI assistance impacts the formation of coding skills. January 29, 2026.
  4. Wolters Kluwer Health. 2026 Future Ready Healthcare Survey Report (with Ipsos). June 2, 2026.
  5. Fernandes, D., Villa, S., Welsch, R., et al. AI makes you smarter but none the wiser: the disconnect between performance and metacognition. Computers in Human Behavior, vol. 175, 108779 (February 2026).
  6. Shaw, S.D. & Nave, G. Thinking — Fast, Slow, and Artificial: How AI Is Reshaping Human Reasoning and the Rise of Cognitive Surrender. Wharton, January 11, 2026.
  7. Koch, C. Beyond the Steeper Curve: AI-Mediated Metacognitive Decoupling and the Limits of the Dunning-Kruger Metaphor. arXiv:2603.29681, March 31, 2026.
  8. Becker, J., Rush, N., Barnes, E., Rein, D. (METR). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. July 10, 2025.
  9. Chan, A. Preference for Explanations: Case of Explainable AI. HBS Working Paper 26-028; summarized in Harvard Business Review, June 2026.
  10. The Oversight Fatigue Problem: Why HITL Breaks Down at Scale. HackerNoon, 2026.
  11. Gartner. 40% of Enterprise Apps Will Feature Task-Specific AI Agents by 2026. August 26, 2025.
  12. Gartner (Daryl Plummer). Top Predictions for IT Organizations and Users in 2026 and Beyond. October 21, 2025.
  13. European Union. EU AI Act, Article 14 (Human Oversight).
  14. European Commission. Draft Commission guidelines on the classification of high-risk AI systems (Article 6). May 19, 2026.
  15. FDA. Clinical Decision Support Software final guidance, January 6, 2026 (see Cooley analysis).
  16. FAA. SAFO 13002 (2013) and SAFO 17007 (2017), Manual Flight Operations.
  17. Ford rehires “gray beard” engineers after AI falls short. TechCrunch, June 28, 2026; see also Bloomberg (June 25, 2026) and Forbes (June 30, 2026).
  18. IBM will hire your entry-level talent in the age of AI. TechCrunch, February 12, 2026.
Issue Credits
Author
Ashwin Telang
Editor-in-Chief
David Lovejoy
Published by Horizon Search Institute, a registered trade name of HSI Research Foundation · EIN 42-1954110 · A Delaware nonprofit corporation · horizonsearch.org