LLM Training Prior Contamination in Compliance Interpretation
LLM Training Prior Contamination in Compliance Interpretation: Failure Modes When Generic Legal Knowledge Overrides Proprietary Organisational Policy
- General-purpose Large Language Models already show a strong tendency to generate legally plausible but factually wrong answers, to accept false legal premises, and to remain overconfident about those answers, which means policy assistants start from a baseline risk of fabricated or misapplied authority before any local-document problem is addedDahl et al. (2024)Intelligence (2024)OpenAI et al. (2024)
- Organisation-specific policy interpretation is especially exposed when the controlling clause is long, exception-heavy, noisy, or poorly placed in context, because models are distractible by irrelevant material, weaker on information in the middle of long context, and still limited at reasoning over retrieved statements even when retrieval is partly successfulChen et al. (2023)Liu et al. (2023)BehnamGhader et al. (2023)Service (2024)
- The most important policy-assistant failure mode is wrong applicability, where the model blends public legal priors, retrieved fragments, and local policy into a coherent answer that sounds authoritative while applying the wrong rule or the wrong level of authorityDahl et al. (2024)Chen et al. (2023)Liu et al. (2023)BehnamGhader et al. (2023)Service (2024)
- Contradictory or stale policy corpora are a separate but interacting source of failure, because policy-coherence work in this repository shows that incoherent policy estates already create unsafe conditions for automated enforcement before Large Language Model synthesis adds another error channelMitchell (2026)
- User review is an unreliable backstop once the model answer appears first, because incorrect Artificial Intelligence support shown before independent judgment lowers human accuracy, participants follow algorithmic recommendations more closely than equally accurate human ones, and even larger recommendation errors are often insufficient to trigger interventionMatute (2024)Jussupow et al. (2024)Goddard et al. (2012)Github (n.d.)
- Policy-assistant outputs can become more dangerous when they increase user confidence rather than raw correctness, because Large Language Model input can more than double human overconfidence and institutionally trusted interfaces can lead users to discount known inaccuracy risksLi et al. (2025)Service (2024)Mitchell (2026)
- Official assurance guidance does not support generic trust in benchmarked model capability as proof of policy safety, because NIST and United Kingdom guidance both frame generative Artificial Intelligence assurance as context-specific evaluation against regulation, standards, limitations, organisational values, and lifecycle monitoring dutiesNational (n.d.)Autio et al. (2024)United (n.d.)United (n.d.)
- The safest architecture for consequential policy interpretation is to use the model only to draft or suggest an interpretation, require citations to the controlling local source, abstain when applicability is not established from authorised documents, and route ambiguous cases into deterministic rules or human escalation rather than letting the model own the final interpretationUnited (n.d.)United (n.d.)United (n.d.)Mitchell (2026)Mitchell (2026)
Research Question
What failure modes emerge when Large Language Models (LLMs) combine generic public legal knowledge with proprietary organisational policy in compliance interpretation tasks?
Findings
(Populated from §6 Synthesis above.)
Executive Summary
In policy interpretation, a Large Language Model can return an authoritative answer that applies the wrong rule or the wrong level of authority.
This risk grows when general-purpose models hallucinate legal authority and when long or weakly retrieved local context gives them multiple ways to miss the controlling internal rule.
This risk also grows when the underlying policy estate is contradictory or stale, because prior repository research shows that incoherent policy sets are already unsafe for automated enforcement before model synthesis adds another error source.
Users are then poorly positioned to catch the mismatch if the model answer appears first, because automation-bias evidence shows that early machine advice reduces human accuracy, human-in-the-loop designs can increase uptake while lowering decision quality, and model input can amplify reviewer overconfidence.
Official assurance guidance therefore supports treating policy assistants as proposal systems that must be tested against local authority boundaries, checked against organisational values and compliance requirements, and paired with explicit abstention and escalation paths for ambiguous cases.
Key Findings
- General-purpose Large Language Models already show a strong tendency to generate legally plausible but factually wrong answers, to accept false legal premises, and to remain overconfident about those answers, which means policy assistants start from a baseline risk of fabricated or misapplied authority before any local-document problem is added.
- Organisation-specific policy interpretation is especially exposed when the controlling clause is long, exception-heavy, noisy, or poorly placed in context, because models are distractible by irrelevant material, weaker on information in the middle of long context, and still limited at reasoning over retrieved statements even when retrieval is partly successful.
- The most important policy-assistant failure mode is wrong applicability, where the model blends public legal priors, retrieved fragments, and local policy into a coherent answer that sounds authoritative while applying the wrong rule or the wrong level of authority.
- Contradictory or stale policy corpora are a separate but interacting source of failure, because policy-coherence work in this repository shows that incoherent policy estates already create unsafe conditions for automated enforcement before Large Language Model synthesis adds another error channel.
- User review is an unreliable backstop once the model answer appears first, because incorrect Artificial Intelligence support shown before independent judgment lowers human accuracy, participants follow algorithmic recommendations more closely than equally accurate human ones, and even larger recommendation errors are often insufficient to trigger intervention.
- Policy-assistant outputs can become more dangerous when they increase user confidence rather than raw correctness, because Large Language Model input can more than double human overconfidence and institutionally trusted interfaces can lead users to discount known inaccuracy risks.
- Official assurance guidance does not support generic trust in benchmarked model capability as proof of policy safety, because NIST and United Kingdom guidance both frame generative Artificial Intelligence assurance as context-specific evaluation against regulation, standards, limitations, organisational values, and lifecycle monitoring duties.
- The safest architecture for consequential policy interpretation is to use the model only to draft or suggest an interpretation, require citations to the controlling local source, abstain when applicability is not established from authorised documents, and route ambiguous cases into deterministic rules or human escalation rather than letting the model own the final interpretation.
Assumptions
- Legal-question answering and grounded public-information chat are used as proxies for enterprise policy assistants because both require selecting the controlling authority from a broader textual environment, but they are not direct measurements of internal corporate policy use.
- Human-review findings from judicial, prediction, and clinical decision-support settings are treated as behavioural proxies for compliance review because the underlying task is evaluation of a machine recommendation under uncertainty.
Analysis
The strongest direct evidence is the combination of legal-domain hallucination studies with context-use studies, because together they explain both why the model can state a wrong rule and why the local controlling rule can fail to displace that error.
One competing explanation is that retrieval quality alone causes the problem. The retriever-augmented reasoning paper and long-context papers make that explanation too narrow, because they show weaknesses both in getting the right evidence and in using it correctly once present.
Another competing explanation is that the policy source itself is incoherent, contradictory, or stale. The policy-coherence item in this repository supports that qualification, and it narrows the central claim here: wrong applicability can arise from weak local policy design alone, and model-plus-context synthesis adds a second error channel on top of that pre-existing condition.
The behavioural evidence matters because an enterprise might otherwise conclude that ordinary reviewer sign-off closes the risk, yet the reviewed studies show the opposite pattern: first-presented machine output can lower human accuracy, and human-in-the-loop workflows can still drift toward rubber-stamping.
For that reason, the evidence weighs in favour of workflow and assurance controls rather than prompt-only fixes, because official guidance repeatedly frames safe use as a matter of context-specific testing, compliance audit, and documented review responsibilities.
Risks, Gaps, and Uncertainties
- Direct public evidence on proprietary internal-policy assistants remains limited, so this synthesis rests on legal-question answering, public-sector policy chat, and broader human-review studies rather than on a large benchmark of enterprise compliance deployments.
- The evidence is stronger on failure mechanisms than on measured mitigation effect sizes for enterprise policy work, because official guidance specifies the control types to use but rarely publishes controlled before-and-after outcome data for organisational deployments.
- The legal-hallucination evidence is strongest for public legal authorities, not for internal corporate standards, so the claim about local-policy applicability remains inferential even though the context-use evidence makes the transfer plausible.
- This item does not quantify which user-interface changes best improve detection of inapplicable answers, because the available sources support independent-first review and verification intensity broadly more strongly than any single interface pattern.
Open Questions
- What benchmark best measures local-policy applicability rather than generic legal correctness?
- How much does mandatory citation of the controlling internal clause improve reviewer detection of context mismatch?
- Which abstention threshold is most practical before a policy assistant should escalate to a deterministic rule or a human specialist?
- How much does chunking or document-layout design change error rates on long policy manuals with exceptions and appendices?
sources
- [x] National Institute of Standards and Technology (NIST) AI Risk Management Framework
- [x] Autio et al. (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- [x] United Kingdom (UK) Government Introduction to Artificial Intelligence Assurance
- [x] United Kingdom (UK) Government Portfolio of Artificial Intelligence Assurance Techniques
- [x] United Kingdom (UK) Government Guidance to Civil Servants on Use of Generative Artificial Intelligence
- [x] United Kingdom (UK) Government Generative Artificial Intelligence Framework for His Majesty's Government (HMG)
- [x] Government Digital Service (2024) The findings of our first generative Artificial Intelligence experiment: GOV.UK Chat
- [x] OpenAI et al. (2024) Generative Pre-trained Transformer 4 (GPT-4) Technical Report
- [x] Dahl et al. (2024) Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
- [x] Stanford Human-Centered Artificial Intelligence (2024) Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive
- [x] Chen et al. (2023) Large Language Models Can Be Easily Distracted by Irrelevant Context
- [x] Liu et al. (2023) Lost in the Middle: How Language Models Use Long Contexts
- [x] BehnamGhader et al. (2023) Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language Model
- [x] Li et al. (2025) Large Language Models are overconfident and amplify human bias
- [x] Goddard et al. (2012) Automation bias: a systematic review of frequency, effect mediators, and mitigators
- [x] Vicente and Matute (2024) The impact of Artificial Intelligence errors in a human-in-the-loop process
- [x] Jussupow et al. (2024) Putting a human in the loop: Increasing uptake, but decreasing accuracy of automated decision-making
- [x] Mitchell (2026) Accountability and governance risks in Artificial Intelligence-assisted policy interpretation
- [x] Mitchell (2026) Pressure for quick closure and confirmation-bias risks in Artificial Intelligence policy interpretation
- [x] Mitchell (2026) Governance Policy Application: Deterministic Requirements vs Stochastic Large Language Model Elements
- [x] Mitchell (2026) Human cognitive bias toward Artificial Intelligence correctness and explainability
- [x] Mitchell (2026) Policy coherence as a machine-checkable prerequisite
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-17 | 49573db | Initial completion |