LLM Response Style and Confidence Signalling

LLM Response Style and Confidence Signalling: How AI Fluency Degrades User Calibration on Uncertainty in Ambiguous Policy Compliance Contexts

2026-05-17 · governance-policy security-risk benchmarks-eval organisational-design · medium · source → · wiki →
key claims
  1. Natural-language uncertainty expressions can materially improve whether user reliance matches answer reliability, because participants exposed to first-person LLM hedging relied on the system less blindly and answered medical questions more accurately than participants shown more assertive wordingKojima et al. (2024)
  2. Explanations and polished answer structures do not automatically create better human-AI decisions, and multiple studies show they can increase acceptance of recommendations even when the underlying recommendation quality does not improveBansal et al. (2021)Schubert et al. (2023)
  3. Information that is easier to process because it is repeated or smoothly presented can raise subjective confidence independently of factual truth, which means a polished policy explanation can feel more reliable than it is before any substantive validation occursStump et al. (2024)
  4. Current LLMs are not dependable narrators of their own certainty, because recent benchmark studies find persistent overconfidence and large gaps between stated confidence and actual correctness even when prompt engineering and consistency checks improve performance somewhatXiong et al. (2024)Shukla et al. (2024)
  5. On legal-style questions that resemble policy interpretation more closely than general trivia does, state-of-the-art LLMs still produce frequent hallucinations, fail to correct false premises reliably, and often need extra probing before their uncertainty becomes visibleDahl et al. (2024)Agrawal et al. (2024)
  6. Ambiguity handling skill does not eliminate interpretation risk, because Kamath et al. report over 90% accuracy in some ambiguity datasets while other studies still show weak matching between stated confidence and actual correctness, plus frequent legal-task hallucinations in real user-facing settingsKamath et al. (2024)Dahl et al. (2024)Xiong et al. (2024)
  7. For policy and compliance assistants, the main governance control should be preserving the user's motivation to verify or escalate, not merely improving prose quality, because wording, aggregation, queue pressure, and the risk of under-reliance all affect whether human review remains meaningfulNational (2023)Schubert et al. (2023)Github (n.d.)Github (n.d.)Dietvorst et al. (2015)

Research Question

How do Large Language Model (LLM) response style and self-reported confidence change how accurately users judge uncertainty and downstream risk when interpreting ambiguous policy and compliance requirements?

Findings

(Populated from §6 Synthesis above.)

Executive Summary

Confident, fluent LLM policy answers are likely to make users underweight ambiguity and overestimate correctness unless the system explicitly signals uncertainty and keeps evidence inspection visible.

The strongest direct evidence is at the user-interface layer: natural-language uncertainty expressions reduce overreliance, while explanations and polished presentation can increase acceptance of advice without improving correctness.

Current LLM self-reported confidence is too weakly matched to actual correctness to substitute for external verification on its own, because recent evaluations show overconfidence and weak error awareness on difficult or legal-style tasks even when some ambiguity-handling capability exists.

For enterprise policy and compliance assistants, a safer default control pattern is uncertainty-forward wording that is matched to evidence and task stakes, direct quotation of governing text, and escalation of ambiguous or high-impact cases rather than a single authoritative answer path.

Key Findings

  1. Natural-language uncertainty expressions can materially improve whether user reliance matches answer reliability, because participants exposed to first-person LLM hedging relied on the system less blindly and answered medical questions more accurately than participants shown more assertive wording.
  2. Explanations and polished answer structures do not automatically create better human-AI decisions, and multiple studies show they can increase acceptance of recommendations even when the underlying recommendation quality does not improve.
  3. Information that is easier to process because it is repeated or smoothly presented can raise subjective confidence independently of factual truth, which means a polished policy explanation can feel more reliable than it is before any substantive validation occurs.
  4. Current LLMs are not dependable narrators of their own certainty, because recent benchmark studies find persistent overconfidence and large gaps between stated confidence and actual correctness even when prompt engineering and consistency checks improve performance somewhat.
  5. On legal-style questions that resemble policy interpretation more closely than general trivia does, state-of-the-art LLMs still produce frequent hallucinations, fail to correct false premises reliably, and often need extra probing before their uncertainty becomes visible.
  6. Ambiguity handling skill does not eliminate interpretation risk, because Kamath et al. report over 90% accuracy in some ambiguity datasets while other studies still show weak matching between stated confidence and actual correctness, plus frequent legal-task hallucinations in real user-facing settings.
  7. For policy and compliance assistants, the main governance control should be preserving the user's motivation to verify or escalate, not merely improving prose quality, because wording, aggregation, queue pressure, and the risk of under-reliance all affect whether human review remains meaningful.

Assumptions

Analysis

The evidence is strongest on two points: users change their verification behavior when wording or interface cues change, and current LLMs remain weak at matching stated confidence to actual correctness.

The main alternative explanation is that better explanations or better model capability could remove the need for explicit uncertainty or escalation controls. The retrieved evidence does not support that stronger claim, because explanations increased acceptance without improving complementary performance, and ambiguity competence did not come with proven confidence-quality matching or error self-awareness.

I therefore weighted evidence about how wording changes user reliance more heavily than raw ambiguity-performance evidence when answering the research question. This does not imply that stronger uncertainty cues should be maximised without limit, because Dietvorst's algorithm-aversion study and later reliance-adjustment work both show that visible error or poorly tuned cues can also produce under-reliance. For governance design, the relevant failure is not only whether the model can sometimes get ambiguity right, but whether users can tell when they should stop trusting the first answer and seek authoritative review without being pushed into blanket disuse.

Risks, Gaps, and Uncertainties

Open Questions


sources

Identified but not consulted

cites
cites How do organisational incentives, culture, and behaviour influence adherence to governance in AI and low-code environments?
cites Human cognitive bias toward Artificial Intelligence (AI) correctness and explainability: automation bias, Reinforcement Learning from Human Feedback (RLHF) sycophancy, and mechanistic interpretability limits
cites How should human-in-the-loop (HITL) design be adapted when AI review volume makes human reviewers a bottleneck or causes rubber-stamping?
cites What capability and control design is needed to mitigate incentive misalignment, shadow Artificial Intelligence (AI), rail bypass, and skill decay at enterprise scale?
cites International Organization for Standardization (ISO) and International Electrotechnical Commission (IEC) 42001:2023 controls, adoption, reputation, and evolution
related (frontmatter)
related When and how should human intervention be incorporated into Artificial Intelligence (AI)-driven and automated workflows?
related Universal Entity Lifecycle Governance Framework (UELGF) extension: human oversight and accountability layer, named owners, escalation paths, and accountability alignment with emerging agentic Artificial Intelligence (AI) governance standards
related Explainable Artificial Intelligence (XAI): current research state, leading institutions, and regulatory intersection in heavily regulated industries
version history
versiondatecommitsummary
1.02026-05-18a2bee71Initial completion

Connected items

Loading…

View full knowledge graph →