Governance Policy Application
Governance Policy Application: Deterministic Requirements vs Stochastic Large Language Model (LLM) Elements
- The NIST Artificial Intelligence Risk Management Framework GOVERN function requires documented legal and regulatory understanding, transparent controls, ongoing monitoring, clear roles, and executive responsibility, which supports treating governance as a documented control process rather than as ad hoc model behavior at the point of decisionNational (n.d.)Iso (n.d.)
- European Commission and Information Commissioner's Office guidance restrict solely automated significant decisions and require meaningful human intervention, contestability, and regular checks, so nominal human involvement around a stochastic model is not enough when the outcome mattersEuropean (n.d.)Information (n.d.)
- Articles 12, 14, and 15 of the European Union Artificial Intelligence Act require lifetime logs, effective human oversight with override or stop capability, and consistent accuracy or robustness, which makes irreproducible free-form model output an inadequate sole control surface for high-risk governance decisionsAI (n.d.)AI (n.d.)AI (n.d.)
- ISO/IEC 38505, ISO/IEC 38507, and Floridi et al. reinforce the same direction by framing acceptable data and Artificial Intelligence use as a governing-body responsibility that must preserve stakeholder confidence, intelligibility, and accountable explanationIso (n.d.)Iso (n.d.)Floridi et al. (2018)
- Microsoft's Azure OpenAI documentation states that repeated calls are nondeterministic by default and that determinism is still not guaranteed even when seed and backend fingerprint controls are held constant, so current provider features improve consistency but do not guarantee reproducibilityMicrosoft (n.d.)
- Thinking Machines Lab's technical analysis indicates that temperature zero does not eliminate practical nondeterminism and links residual variance to serving-path behavior such as batch-sensitive execution, so decoding settings alone are not a sufficient governance controlLab (2025)
- Governance decisions that approve, deny, classify, sanction, or trigger reporting or side effects therefore need deterministic rules, logged thresholds, or empowered human adjudication as the final authority, because that is where traceability and contestability are tested in practiceEuropean (n.d.)Information (n.d.)AI (n.d.)Compliance (n.d.)
- Controlled probabilistic variation is acceptable for preparatory or assistive tasks only when outputs are logged and routed through deterministic or meaningfully supervised final gates before any consequential action is executedNational (n.d.)AI (n.d.)AI (n.d.)Github (n.d.)
Research Question
To what extent must governance policy application be deterministic, consistent, reproducible, and auditable, versus allowing stochastic or probabilistic elements when Artificial Intelligence (AI) or Large Language Models (LLMs) are involved?
Findings
Executive Summary
Governance policy application must be deterministic at the final decision surface whenever an output can create legal, rights-significant, compliance, security, or hard-to-reverse operational effects, because the strongest official sources require traceability, meaningful oversight, consistent performance, and contestable outcomes. Stochastic Large Language Model behavior is acceptable upstream in assistive tasks such as summarization, option generation, and draft rationale production, but only when logging and human or rule-based final authority remain in place before action. Current deployed-model controls do not make Large Language Model output fully reproducible, because Microsoft documents residual nondeterminism even with reproducibility features and engineering analysis attributes remaining variance to inference-serving behavior rather than sampling alone. The practical design rule is therefore to place determinism in the final control gate, the evidence record, and the meaningful oversight path instead of expecting the model itself to satisfy governance-grade reproducibility.
Key Findings
- The NIST Artificial Intelligence Risk Management Framework GOVERN function requires documented legal and regulatory understanding, transparent controls, ongoing monitoring, clear roles, and executive responsibility, which supports treating governance as a documented control process rather than as ad hoc model behavior at the point of decision.
- European Commission and Information Commissioner's Office guidance restrict solely automated significant decisions and require meaningful human intervention, contestability, and regular checks, so nominal human involvement around a stochastic model is not enough when the outcome matters.
- Articles 12, 14, and 15 of the European Union Artificial Intelligence Act require lifetime logs, effective human oversight with override or stop capability, and consistent accuracy or robustness, which makes irreproducible free-form model output an inadequate sole control surface for high-risk governance decisions.
- ISO/IEC 38505, ISO/IEC 38507, and Floridi et al. reinforce the same direction by framing acceptable data and Artificial Intelligence use as a governing-body responsibility that must preserve stakeholder confidence, intelligibility, and accountable explanation.
- Microsoft's Azure OpenAI documentation states that repeated calls are nondeterministic by default and that determinism is still not guaranteed even when seed and backend fingerprint controls are held constant, so current provider features improve consistency but do not guarantee reproducibility.
- Thinking Machines Lab's technical analysis indicates that temperature zero does not eliminate practical nondeterminism and links residual variance to serving-path behavior such as batch-sensitive execution, so decoding settings alone are not a sufficient governance control.
- Governance decisions that approve, deny, classify, sanction, or trigger reporting or side effects therefore need deterministic rules, logged thresholds, or empowered human adjudication as the final authority, because that is where traceability and contestability are tested in practice.
- Controlled probabilistic variation is acceptable for preparatory or assistive tasks only when outputs are logged and routed through deterministic or meaningfully supervised final gates before any consequential action is executed.
Assumptions
- This item treats internal governance decisions that can change access, enforcement, escalation, or reporting state as materially similar to rights-significant automated decisions even when Article 22 may not formally apply, because the same traceability and contestability logic still shapes defensible governance design.
- A deterministic final decision can be implemented either through encoded policy logic or through a human reviewer with authority, evidence, and override capability, because the reviewed sources require effective oversight and accountability but do not prescribe one universal architecture.
Analysis
The evidence weights official governance and regulatory sources most heavily, because they define what a defensible policy-application process must contain even when they do not prescribe one vendor or architecture. That weighting makes the core conclusion procedural rather than philosophical: determinism is required where an organization commits to an outcome, not necessarily where a model explores options or drafts reasoning upstream. One plausible rival view is that better prompting, lower temperature, or seeded sampling can make the model deterministic enough, but Microsoft's own documentation rejects guaranteed determinism and the technical analysis explains why residual variance persists after those controls. Another rival view is that blanket human review can compensate for model variability, but the Information Commissioner's Office and the European Union Artificial Intelligence Act both imply that oversight must be meaningful, empowered, and resistant to automation bias rather than a fast approval queue. The best-supported operating model is therefore a tiered one: stochastic systems can help interpret, search, summarize, or draft, but deterministic rules, explicit evidence capture, and empowered human or policy authority must own the final governance commitment.
Risks, Gaps, and Uncertainties
- The accessible ISO sources are official summaries rather than full standard text, so the ISO-based portion of the argument is less clause-specific than the NIST, European Commission, Information Commissioner's Office, and European Union Artificial Intelligence Act portions.
- The European Union Artificial Intelligence Act sources directly govern high-risk systems, so extending their operational logic to lower-risk internal governance workflows is a reasoned design inference rather than a direct legal holding for every use case.
- Public reproducibility documentation demonstrates residual variability, but it does not provide a comprehensive cross-provider benchmark for how much variance remains across prompt classes, model families, or runtime conditions.
Open Questions
- What minimum joined log schema is sufficient to reconstruct a mixed model-and-policy governance decision across multiple vendors and workflow engines?
- In which non-high-risk but still material governance workflows should double human verification be required, rather than a single empowered reviewer or deterministic rule engine?
- How should organizations set escalation thresholds for converting Large Language Model recommendations into automatically accepted actions without reintroducing symbolic review?
sources
- [x] National Institute of Standards and Technology (NIST) Artificial Intelligence Risk Management Framework Core
- [x] National Institute of Standards and Technology (NIST) (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- [x] European Commission Restrictions on automated decision-making
- [x] Information Commissioner's Office What is the impact of Article 22 of the United Kingdom (UK) General Data Protection Regulation (GDPR) on fairness?
- [x] AI Act Service Desk Article 12 Record-keeping
- [x] AI Act Service Desk Article 14 Human oversight
- [x] AI Act Service Desk Article 15 Accuracy, robustness and cybersecurity
- [x] ISO/IEC 38505-1:2017 Information technology, Governance of IT, Governance of data
- [x] ISO/IEC 38507:2022 Information technology, Governance implications of the use of Artificial Intelligence by organizations
- [x] Floridi et al. (2018) AI4People, An Ethical Framework for a Good AI Society
- [x] Microsoft Learn How to generate reproducible output with Azure OpenAI
- [x] Thinking Machines Lab (2025) Defeating Nondeterminism in LLM Inference
- [x] Hybrid Architecture Design: Probabilistic Large Language Models for Interpretation, Deterministic Layers for Governance Enforcement
- [x] Compliance Risks of Relying on Stochastic Large Language Model Outputs for Governance, Privacy, and Regulatory Decisions
- [x] Data Governance Standards and Regulations Applied to Artificial Intelligence Systems and Multi-Step Autonomous AI Deployments
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-10 | 7b9b4fc | Initial completion |