AI inverted the knowledge-work scarcity equation

AI inverted the knowledge-work scarcity equation: volume is free, correctness is the scarce resource

2026-03-14 · agentic-ai workforce-skills cost-performance · medium · source → · wiki →
key claims
  1. Volume up, delivery flat: AI coding assistants increased individual task completion by 21% and PR merges by 98% across 10,000+ developers in the Faros AI 2025 telemetry study, yet organisational delivery metrics — throughput, stability, change failure rate — remained flat
  2. 3x top-decile quality with structured AI use: In the Dell'Acqua et al. (2025) pre-registered RCT at P&G, teams using AI were 9.2 percentage points more likely to produce solutions rated in the top decile by expert judges, compared to a control mean of 5.8%, corresponding to approximately three times more chances of reaching top-decile quality
  3. AI substitutes for average collaboration, not peak collaboration: AI-enabled individuals in the P&G RCT matched the average quality of two-person human teams working without AI, but the highest quality outputs still came from human teams augmented by AI; peak collaborative judgment, not average competence, is the differentiating constraint
  4. The verification bottleneck is empirically measured: Faros AI 2025 data shows PR review time increased 91% and PR size grew 154% alongside high AI adoption; the human review step did not accelerate to match AI generation speed
  5. The agentic tarpit names the multi-agent correctness failure mode: Wes McKinney's 2025 blog post "The Mythical Agent-Month" identified that parallel AI agent sessions produce contradictory and bloated outputs at machine speed, requiring human reconciliation that cannot be accelerated by adding more agents, creating a hard bottleneck at human conceptual integrity
  6. Workslop is a measurable organisational consequence of volume without correctness discipline: BetterUp Labs (2025) research found approximately 40% of workers received poor-quality AI-generated content in the past month, estimated at 15% of total workplace content, generating coordination overhead and eroding interpersonal trust
  7. The prototype-to-production gap is structural and unresolved: AI has dramatically compressed the time from blank slate to working proof of concept, but judgment-intensive production work — requirements translation, architectural trade-offs, distributed system debugging, and production-readiness validation — remains human-bottlenecked and has not been equivalently accelerated by current AI tooling
  8. AI breaks functional silos in innovation tasks: The P&G RCT found that AI-assisted Research and Development (R&D) and commercial professionals independently produced balanced, integrated proposals regardless of professional background, demonstrating that AI can reduce quality losses from siloed domain expertise by giving individuals access to cross-functional competency

Research Question

Before Artificial Intelligence (AI), throughput (volume of output) was the binding constraint on knowledge work. AI has dramatically reduced the cost of generating output. Does the evidence support the thesis that correctness — whether a given output is architecturally sound, strategically coherent, and fit for production — is now the primary scarce resource? And if so, what practices and team structures actually increase correctness rather than merely volume?

Specifically: how does shared mental model size (a function of team size) interact with the volume of AI-generated output to determine a team's real productivity — and what does this imply for how quality should be managed in the AI era?

Findings

Executive Summary

Correctness — whether output is production-ready, strategically coherent, or factually accurate — has become the primary constraint in AI-augmented knowledge work; volume generation is no longer the limiting factor. Faros AI's 2025 telemetry study of 10,000+ developers recorded individual task completion up 21% and PR merges up 98%, yet organisational delivery metrics remained flat — the "AI Productivity Paradox." The Dell'Acqua et al. (2025) pre-registered RCT at P&G explains why: AI-assisted teams were approximately three times more likely to produce top-10% quality outputs, but only when human judgment was structurally embedded through training, guided prompting, and expert evaluation; without that structure, AI generates volume at machine speed while verification remains bottlenecked by human attention. Wes McKinney named this the "agentic tarpit" — parallel AI sessions produce contradictory, bloated outputs faster than human judgment can triage them — and industry researchers documented its organisational expression as "workslop": AI-generated content that looks professional but lacks substance, received by approximately 40% of workers in the past month. Increasing correctness requires the same disciplined practices that always mattered — small batches, explicit quality standards, domain expertise — which organisations that master them can use to differentiate when competitors are focused on volume. [inference — derived from evidence convergence; the competitive advantage framing is interpretive]

Key Findings

  1. Volume up, delivery flat: AI coding assistants increased individual task completion by 21% and PR merges by 98% across 10,000+ developers in the Faros AI 2025 telemetry study, yet organisational delivery metrics — throughput, stability, change failure rate — remained flat. [high confidence]

  2. 3x top-decile quality with structured AI use: In the Dell'Acqua et al. (2025) pre-registered RCT at P&G, teams using AI were 9.2 percentage points more likely to produce solutions rated in the top decile by expert judges, compared to a control mean of 5.8%, corresponding to approximately three times more chances of reaching top-decile quality. [high confidence]

  3. AI substitutes for average collaboration, not peak collaboration: AI-enabled individuals in the P&G RCT matched the average quality of two-person human teams working without AI, but the highest quality outputs still came from human teams augmented by AI; peak collaborative judgment, not average competence, is the differentiating constraint. [high confidence]

  4. The verification bottleneck is empirically measured: Faros AI 2025 data shows PR review time increased 91% and PR size grew 154% alongside high AI adoption; the human review step did not accelerate to match AI generation speed. [high confidence]

  5. The agentic tarpit names the multi-agent correctness failure mode: Wes McKinney's 2025 blog post "The Mythical Agent-Month" identified that parallel AI agent sessions produce contradictory and bloated outputs at machine speed, requiring human reconciliation that cannot be accelerated by adding more agents, creating a hard bottleneck at human conceptual integrity. [high confidence]

  6. Workslop is a measurable organisational consequence of volume without correctness discipline: BetterUp Labs (2025) research found approximately 40% of workers received poor-quality AI-generated content in the past month, estimated at 15% of total workplace content, generating coordination overhead and eroding interpersonal trust. [medium confidence]

  7. The prototype-to-production gap is structural and unresolved: AI has dramatically compressed the time from blank slate to working proof of concept, but judgment-intensive production work — requirements translation, architectural trade-offs, distributed system debugging, and production-readiness validation — remains human-bottlenecked and has not been equivalently accelerated by current AI tooling. [high confidence]

  8. AI breaks functional silos in innovation tasks: The P&G RCT found that AI-assisted Research and Development (R&D) and commercial professionals independently produced balanced, integrated proposals regardless of professional background, demonstrating that AI can reduce quality losses from siloed domain expertise by giving individuals access to cross-functional competency. [high confidence]

  9. No universal cross-domain correctness index exists: A search across empirical literature and industry practice found no standardised measurement framework for correctness applicable across code, strategy, and content domains; organisations lack the tooling to measure the constraint they must now manage most urgently. [medium confidence]

  10. Small batches and explicit quality standards are the evidence-backed correctness practices: DORA 2025 identifies working in small batches as the practice most reliably associated with amplifying AI's positive effects; Shopify CEO Toby Lütke's 2025 AI mandate frames explicit taste standards and proof-of-concept discipline as the structural complement to AI-generated output, but no systematic effectiveness study has been published on either practice in the AI context. [medium confidence]

Assumptions

  1. The P&G RCT findings (3x top-decile quality) generalise directionally to other knowledge-work domains, though the specific magnitude applies only to consumer goods innovation tasks structured under experimental conditions. Justification: The mechanism (structured AI access + domain expertise + evaluation framework) is domain-agnostic; the magnitude is domain-specific.

  2. "Correctness" is interpreted as "fitness for purpose" in the relevant domain — expert-judged quality for innovation, bug-free production-ready delivery for software, accuracy + utility for content — not as formal mathematical correctness. Justification: No domain-agnostic definition of correctness was found; all empirical measures are domain-specific operationalisations.

  3. The Faros AI sample (mid-to-large engineering organisations) may over-represent enterprise software teams. AI-native startups with different team structures and higher domain expertise per capita may show different patterns. Justification: No AI-native startup telemetry study of comparable scale was found.

Analysis

AI accelerates the generation phase of knowledge work without equivalently accelerating the verification phase. [inference — derived from Faros AI telemetry showing flat DORA metrics despite volume increases, and from the P&G RCT showing quality improvements only when verification scaffolding is present] Because verification is the binding constraint, organisational throughput does not improve commensurately with volume. [inference — derived from the same evidence] The P&G RCT establishes this experimentally; Faros AI and DORA data measure it at scale; McKinney's agentic tarpit explains the mechanism.

A critical distinction emerges from the P&G study: AI substitutes for average peer collaboration, not peak collaborative quality. [inference — derived from Dell'Acqua et al. 2025 treatment arm comparisons] Solo practitioners using AI can compete with unaugmented teams on typical tasks, meaning the baseline of acceptable output has been raised for everyone. Organisations and teams that differentiate on quality are those that combine strong human judgment with AI — not those that substitute AI for human judgment. [inference — extrapolated from P&G study design; not directly tested as an organisational strategy]

The competing interpretation — that flat organisational metrics reflect adoption friction rather than a structural constraint — is not supported by the evidence. DORA 2024 data shows degradation under high adoption, not improvement. The verification bottleneck is cognitive, not technical: no amount of faster AI makes domain expertise faster to acquire or contextual judgment faster to exercise. [inference — derived from evidence; the cognitive-bottleneck characterisation is analytical, not empirically measured]

Investing in human judgment capacity is the higher-leverage response to the volume-correctness inversion — particularly the judgment required for verification, architectural decision-making, and domain expertise. [inference — derived from evidence synthesis; prescriptive framing is the authors' interpretation, not a finding from any single study] Practices, hiring, and tooling calibrated to this constraint will be more effective than those calibrated to maximising AI-generated volume.

Risks, Gaps, and Uncertainties

Open Questions

  1. Can AI-assisted verification tools (automated review, AI code auditors, fact-checking agents) close the verification bottleneck, or is the bottleneck fundamentally cognitive and therefore not addressable by adding more AI? If the latter, what is the correct investment model?
  2. How should organisations operationalise "correctness" measurement in strategy and content domains at scale, given the absence of a standardised framework?
  3. What is the correct ratio of AI-output volume to human-review capacity, and how does this ratio change as AI model quality improves?
  4. Does the P&G finding (non-core employees reaching expert-quality outputs with AI) imply that small teams can maintain correctness standards on cross-functional tasks by relying on AI for the competence they lack? If so, what are the failure modes when that reliance exceeds a threshold?

Output

sources


Connected items

Loading…

View full knowledge graph →