Extending Traditional Data Governance Frameworks to Address Large Language…
Extending Traditional Data Governance Frameworks to Address Large Language Model (LLM) Non-Determinism and Uncertainty About Deployed Behavior
- DAMA-DMBOK, ISO/IEC 38505, and COBIT remain usable baseline frameworks for deployed Large Language Model governance because they already define stewardship, accountability, metadata, quality, security, and monitoring domains, but they require explicit reinterpretation for generative-system control surfacesDAMA (n.d.)Iso (n.d.)ISACA (2025)
- Metadata and lineage governance must expand from business-data catalogs and pipeline lineage to versioned prompt templates, system instructions, model identifiers, evaluation sets, transparency artifacts, impact assessments, and generation provenance if organizations want auditable Large Language Model operationsDAMA (n.d.)Google (n.d.)Google (n.d.)Microsoft (2022)NIST (2024)
- Classical data-quality governance is insufficient on its own, because deployed generative systems also require behavioral evaluation, confabulation tracking, red teaming, and release criteria that measure whether outputs remain fit for purpose under realistic and adversarial conditionsMicrosoft (2022)Google (n.d.)NIST (2024)
- Governance of prompt templates is a first-class extension point for traditional governance frameworks, because system-level policies, prompt templates, safeguard thresholds, and tool-use constraints materially change model behavior and therefore need documented ownership, review, and revision historyGoogle (n.d.)Google (n.d.)Google (n.d.)Google (n.d.)
- Accountability and oversight domains must be extended so that Large Language Models operate as bounded proposal or interpretation layers while deterministic policy, approval, rollback, and human-override mechanisms retain final authority for consequential governance actionsNIST (n.d.)Microsoft (2022)Governance Policy Application (n.d.)Hybrid Architecture Design (n.d.)Compliance (n.d.)
- The currently lower-friction extension strategy is a mapped control stack in which traditional governance domains absorb generative-AI-specific practices such as content provenance, transparency notes, incident disclosure, safeguards, and ongoing assurance review, even though a separate framework could still offer clearer clause-level guidance in some sectorsDAMA (n.d.)Iso (n.d.)ISACA (2025)NIST (2024)Microsoft (2022)Google (n.d.)
Research Question
How can traditional data governance frameworks be extended or mapped to address the inherent non-determinism and uncertainty about whether deployed behavior remains aligned with intended use in modern Large Language Models (LLMs) and multi-step agent systems?
Findings
Executive Summary
Traditional data governance frameworks can be extended for Large Language Model systems by treating prompt templates, model versions, evaluation evidence, and inference provenance as governed assets and by keeping final consequential authority outside stochastic model output. The accessible public material for DAMA-DMBOK, ISO/IEC 38505, and COBIT already covers accountability, stewardship, metadata, quality, security, and monitoring, but it does not specify how to govern probabilistic outputs, prompt changes, or model-version drift in deployed generative systems. NIST's generative profile plus Microsoft's and Google's responsible-AI frameworks supply the missing operational extensions: impact assessments, intended-use restrictions, content provenance, transparency artifacts, safeguards, red teaming, ongoing evaluation, and incident handling. Alignment uncertainty should be governed as a continuous assurance and change-control problem rather than folded into classical data quality alone, because behavior can shift through prompt interaction, safeguard tuning, and model or backend changes even when business data remain stable.
Key Findings
- DAMA-DMBOK, ISO/IEC 38505, and COBIT remain usable baseline frameworks for deployed Large Language Model governance because they already define stewardship, accountability, metadata, quality, security, and monitoring domains, but they require explicit reinterpretation for generative-system control surfaces.
- Metadata and lineage governance must expand from business-data catalogs and pipeline lineage to versioned prompt templates, system instructions, model identifiers, evaluation sets, transparency artifacts, impact assessments, and generation provenance if organizations want auditable Large Language Model operations.
- Classical data-quality governance is insufficient on its own, because deployed generative systems also require behavioral evaluation, confabulation tracking, red teaming, and release criteria that measure whether outputs remain fit for purpose under realistic and adversarial conditions.
- Governance of prompt templates is a first-class extension point for traditional governance frameworks, because system-level policies, prompt templates, safeguard thresholds, and tool-use constraints materially change model behavior and therefore need documented ownership, review, and revision history.
- Accountability and oversight domains must be extended so that Large Language Models operate as bounded proposal or interpretation layers while deterministic policy, approval, rollback, and human-override mechanisms retain final authority for consequential governance actions.
- The currently lower-friction extension strategy is a mapped control stack in which traditional governance domains absorb generative-AI-specific practices such as content provenance, transparency notes, incident disclosure, safeguards, and ongoing assurance review, even though a separate framework could still offer clearer clause-level guidance in some sectors.
Assumptions
- Public summaries of DAMA-DMBOK, ISO/IEC 38505, and COBIT are sufficiently representative to support domain-level mapping even though the full standards texts are partly paywalled.
- Prompt templates, system instructions, and safeguard policies can be treated as governed metadata assets even when older framework wording does not name them explicitly, because public vendor guidance already treats them as documented and reviewable operational artifacts.
Analysis
The evidence favors extension by mapping rather than replacement because the older frameworks still describe the right governance categories, but newer sources add the operational evidence required for deployed generative systems. The strongest cross-source convergence sits around four extensions: provenance, evaluation, safeguards, and accountability, because NIST, Microsoft, and Google each publish controls in those areas even though they differ in terminology. The practical mapping is therefore: governance and accountability to impact assessment and role separation; metadata and lineage to prompts, models, and generation provenance; quality to behavioral evaluation and confabulation thresholds; security to safeguards and prompt-injection defenses; and audit to transparency notes, release evidence, and incident disclosure. An alternative interpretation is that generative-AI governance now needs a separate framework because legacy standards are summary-level and partly paywalled, and NIST's generative profile already behaves like a specialized companion framework. The evidence still favors mapped extension as the lower-friction adoption path, because DAMA-DMBOK, ISO/IEC 38505, and COBIT continue to provide the organizational ownership and governance categories while NIST and vendor frameworks supply the missing operational detail. Prior repository items sharpen the final control recommendation by adding empirical and governance-specific evidence that current Large Language Model systems remain only partially reproducible, which makes deterministic external authority the safer interpretation of traditional accountability obligations.
Risks, Gaps, and Uncertainties
- Public access to DAMA-DMBOK, COBIT, and ISO/IEC 38505 is summary-level, so some domain mappings are stronger at the category level than at the clause level.
- Google and Microsoft provide implementation-rich guidance, but those sources are vendor frameworks rather than neutral cross-industry standards, so they strengthen the operational extension pattern more than they settle sector-independent minimum requirements.
- The public evidence is stronger for inference-time governance and behavioral assurance than for formal amendment text inside legacy frameworks themselves, which means the mapped extension is better supported than any claim that the older frameworks have already been rewritten comprehensively for Large Language Models.
Open Questions
- Which minimum inference-log fields should become a de facto standard for cross-vendor Large Language Model governance, especially where providers expose different model-version and backend metadata?
- How should organizations set escalation thresholds when evaluation results are mixed across adversarial, safety, and fit-for-purpose benchmarks and behavior may drift away from intended use?
- What governance pattern best handles multi-step agent workflows that chain multiple models, tools, and safeguard layers across organizational boundaries?
sources
- [x] DAMA International DAMA-DMBOK Data Management Body of Knowledge
- [x] DAMA DMBOK 3.0 Project
- [x] ISO/IEC 38505-1:2017 Information technology, Governance of IT, Governance of data
- [x] ISACA COBIT Resource Center
- [x] ISACA (2025) Leveraging COBIT for Effective AI System Governance
- [x] NIST AI Risk Management Framework Overview
- [x] NIST AI Risk Management Framework Core
- [x] NIST AI Risk Management Framework Playbook
- [x] NIST (2024) Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- [x] Microsoft (2022) Responsible AI Standard v2 General Requirements
- [x] Microsoft (2022) Responsible AI Impact Assessment Template
- [x] Google AI Responsible Generative AI Toolkit Overview
- [x] Google AI Responsible Generative AI Toolkit Design a Responsible Approach
- [x] Google AI Responsible Generative AI Toolkit Evaluation
- [x] Google AI Responsible Generative AI Toolkit Safeguards
- [x] Google AI Model Alignment
- [x] Arrieta et al. (2020) Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI
- [x] Data Governance Standards and Regulations Applied to Artificial Intelligence Systems and Multi-Step Autonomous AI Deployments
- [x] Governance Policy Application: Deterministic Requirements vs Stochastic Large Language Model Elements
- [x] Hybrid Architecture Design: Probabilistic Large Language Models for Interpretation, Deterministic Layers for Governance Enforcement
- [x] Implementation Patterns for Regulatory Compliance in Artificial Intelligence-Driven Data Governance: Policy-as-Code, Guardrails, and Output Validation
- [x] Practical Limits of Large Language Model Determinism: Temperature Zero, Fixed Seeds, and Constrained Prompts
- [x] Compliance Risks of Relying on Stochastic Large Language Model Outputs for Governance, Privacy, and Regulatory Decisions
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-11 | d2cdfc2 | Initial completion |