Updating the enterprise Artificial Intelligence ecosystem capability reference…
Updating the enterprise Artificial Intelligence ecosystem capability reference architecture using second-cycle 2026-05 completed items
- The existing five-layer architecture remains usable, but the provenance plane now has to cover legibility, meaning catalog, configuration, topology, and drift visibility, and runtime evidence in addition to design-time lineage if it is to explain what the system actually became in operationBackstage (n.d.)ServiceNow (n.d.)Dynatrace (n.d.)Mitchell (2026)Mitchell (2026)
- Delivery and platform engineering now need an explicit declared AIBOM capability, where AIBOM is an AI bill-of-materials artifact for models, data, configuration, prompts, tools, and related execution context, because both managed platforms and code-native orchestration frameworks require deliberate extraction paths before approved design can be compared with later runtime behaviorOwaspaibom (n.d.)CycloneDX (n.d.)Mitchell (2026)Mitchell (2026)Mitchell (2026)
- Runtime-observed evidence has to become a first-class architecture service, because prompts, retrieved context, tool order, authority handoffs, and missing observability are governance-relevant facts that only appear after execution beginsMitchell (2026)Mitchell (2026)Mitchell (2026)
- The orchestration layer now needs explicit delegation-aware identity and tool-action components, because portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receiptMitchell (2026)Mitchell (2026)Mitchell (2026)
- Operating assurance now needs named architecture components for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weakMitchell (2026)Mitchell (2026)Mitchell (2026)
- Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architectureMitchell (2026)Mitchell (2026)Mitchell (2026)
- Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layersMitchell (2026)Github (n.d.)Github (n.d.)
- Central ownership should stay with provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds tied to local business contextMitchell (2026)Mitchell (2026)Mitchell (2026)
Research Question
How should the enterprise Artificial Intelligence (AI) ecosystem capability reference architecture (as expressed in 2026-04-22-enterprise-ai-capability-model, 2026-05-05-enterprise-ai-capability-stack, and 2026-05-06-ai-capability-reference-architecture-security-supply-chain-update) be revised and extended to incorporate findings from the second-cycle 2026-05 completed items, covering AI Bill of Materials (AIBOM) declared construction practices, effectiveness and risk-mitigation limits, multi-agent identity attribution, platform observability controls, European Union (EU) AI Act regulatory intersection, and OpenTelemetry-based runtime capture; AI production incidents; regulatory guidance updates; Five Eyes AI risk posture; open-weight model safeguard policies; and Information Technology (IT) legibility and measurement frameworks, that were not incorporated in the first-cycle update?
Findings
(Populated from §6 Synthesis above.)
Executive Summary
The enterprise AI reference architecture should remain a five-layer model, but the shared provenance and governance planes now need explicit second-cycle services for legibility, meaning catalog, configuration, topology, and drift visibility, runtime evidence, operational assurance, and regulator-facing governance.
Second-cycle evidence shows that design-time inventory alone is insufficient, because the architecture also has to preserve runtime divergence, delegation receipts, incident-triggering conditions, and estate legibility across catalogs, topology, and structural-drift evidence.
Regulatory and Five Eyes updates also move operating-model work into the architecture itself, because board literacy, supplier-risk management, secure logging, rollback authority, and evidence assembly are now observable control surfaces instead of background management assumptions.
The practical revision is to widen the provenance plane into a provenance-and-legibility plane, widen the governance plane into a governance, evaluation, and operational-assurance plane, and keep shared semantics and retained evidence under explicit central ownership.
Key Findings
- The existing five-layer architecture remains usable, but the provenance plane now has to cover legibility, meaning catalog, configuration, topology, and drift visibility, and runtime evidence in addition to design-time lineage if it is to explain what the system actually became in operation.
- Delivery and platform engineering now need an explicit declared AIBOM capability, where AIBOM is an AI bill-of-materials artifact for models, data, configuration, prompts, tools, and related execution context, because both managed platforms and code-native orchestration frameworks require deliberate extraction paths before approved design can be compared with later runtime behavior.
- Runtime-observed evidence has to become a first-class architecture service, because prompts, retrieved context, tool order, authority handoffs, and missing observability are governance-relevant facts that only appear after execution begins.
- The orchestration layer now needs explicit delegation-aware identity and tool-action components, because portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receipt.
- Operating assurance now needs named architecture components for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weak.
- Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architecture.
- Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layers.
- Central ownership should stay with provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds tied to local business context.
Assumptions
- Assumption: Enterprises can emit enough runtime telemetry to reconcile approved design and observed execution at least for high-risk workflows. Justification: the runtime-capture and platform-observability items treat telemetry enablement as difficult but operationally feasible.
- Assumption: Financial-services and security-sensitive enterprises will continue to prefer centralized evidence retention and reporting even when execution remains federated. Justification: the regulatory-delta and Five Eyes items both point toward stronger centralized accountability rather than looser local proof models.
- Assumption: Open-weight safeguard services will stay optional because their infrastructure cost and policy-design overhead will not be justified for every workflow. Justification: the safeguard item supports selective, second-stage use rather than universal inline placement.
Analysis
The second-cycle evidence makes the architecture more operational, not more abstract, because the missing pieces are evidence-handling and control-surface capabilities that only appear when systems are released, observed, challenged, and regulated in practice.
The main trade-off is between central coherence and layer-local enforcement. Centralizing everything would hide where trust decisions really occur, while leaving every team to invent its own evidence model would destroy comparability, auditability, and incident reconstruction.
The best-supported resolution is to keep control semantics and retained evidence centralized while leaving enforcement near the layer that owns the relevant trust decision, retrieval permissions in data and knowledge, semantic safeguards near inference, tool and delegation controls in orchestration, and release-time signing, extraction, and evaluation in delivery and platform engineering.
A major alternative would be to keep every second-cycle addition inside the existing layers without widening either cross-cutting plane, or to replace the baseline stack with a flatter capability mesh, but the evidence weighs against both moves because the new requirements cluster around shared evidence, legibility, and retained-governance services that cut across every layer rather than belonging cleanly to one local component family.
Plausible alternative placements exist for open-weight safeguard capability, especially model-layer controls and pre-deployment review gates, but the evidence supports locating the canonical policy service in the governance and evaluation plane so that delivery gates and runtime checkpoints can invoke one shared policy logic rather than duplicate policy semantics in each layer.
Revised component list:
- Data and knowledge layer: authoritative-source registry, permission-safe retrieval, retrieval-snapshot metadata, and upstream provider-disclosure links.
- Model and inference layer: approved model and provider registry, semantic safeguards, policy-conditioned classification hooks, and model-facing runtime signal capture.
- Orchestration and execution layer: tool allowlists, delegation-receipt capture, action checkpoints, recursion and stop authority, and per-run authority context.
- Delivery and platform-engineering layer: declared AIBOM builder, signed artifact lineage, promotion-time evaluation gates, telemetry collector pipeline, and approved-versus-observed divergence classifier.
- Operating-model layer: board and executive literacy, supplier and concentration-risk governance, fairness review, incident response, rollback authority, and regulator-facing evidence assembly.
- Provenance-and-legibility plane: declared and observed supply-chain records, ownership-bearing catalogs, architecture blueprints, runtime topology, and structural-drift evidence.
- Governance, evaluation, and operational-assurance plane: policy semantics, evaluation standards, exceptions, explainability artifacts, incident records, and optional second-stage safeguard review.
Ownership recommendations:
- Central platform, security, and risk functions should own provenance schema, telemetry normalization, retained evidence, evaluation standards, and regulatory reporting because those assets lose value when fragmented by domain.
- Domain teams should own authoritative content, workflow composition, business-specific thresholds, and local exception context because those choices depend on business meaning and risk appetite that the central platform cannot infer.
- Shared services should expose signed interfaces and evidence contracts rather than monolithic approval queues, because second-cycle evidence favors strong shared semantics with layer-local enforcement and action-specific stop authority.
Risks, Gaps, and Uncertainties
- Runtime-evidence quality remains sensitive to platform configuration and adapter quality, so an architecture can still overstate observability if collectors, spans, or export paths are only partially enabled.
- The architecture can reduce structural and operational blind spots without eliminating semantic failure modes, because AIBOM and telemetry remain weaker than adversarial content and authority misuse at proving behavior is safe.
- The strongest financial-services regulatory evidence in this cycle comes from APRA and cross-jurisdiction synthesis rather than from a uniform new standard across all reviewed regulators, so some operating-model recommendations remain medium-confidence generalizations.
- Legibility metrics are still composite rather than standardized, so estates may need local scorecards before cross-domain comparisons become reliably meaningful.
Open Questions
- What minimum delegation-receipt schema would let orchestration runtimes preserve subject, actor, scope, target, and approval context across cross-platform tool calls?
- Which minimal composite legibility dashboard can compare catalog coverage, runtime topology coverage, divergence rates, and structural-drift findings without becoming another stale governance artifact?
- When does a policy-conditioned safeguard service justify its infrastructure cost compared with deterministic rules plus sampled human review?
- Which evidence thresholds should trigger automatic rollback versus manual escalation when approved design and observed runtime diverge in high-risk workflows?
sources
- [x] Mitchell (2026) Enterprise AI capability model for use-case maturity decisions: baseline five-layer reference architecture
- [x] Mitchell (2026) Enterprise capability stack for sustainable multi-provider Artificial Intelligence: shared-control-core synthesis that the architecture extends
- [x] Mitchell (2026) Integrating 2026-05 security and supply chain findings into the enterprise AI capability reference architecture: first-cycle update; establishes the immediate baseline for this second-cycle revision
- [x] Mitchell (2026) AIBOM declared construction practice: operational practice for constructing declared AIBOM artifacts
- [x] Mitchell (2026) AIBOM effectiveness and risk-mitigation limits: where AIBOM controls are and are not effective
- [x] Mitchell (2026) AIBOM identity attribution in multi-agent systems: practice: multi-agent identity attribution practices and architectural implications
- [x] Mitchell (2026) AIBOM platform observability and control comparison: comparative view of platform observability controls for AIBOM
- [x] Mitchell (2026) AIBOM and EU AI Act regulatory intersection: compliance obligations and intersection between AIBOM and the European Union (EU) AI Act
- [x] Mitchell (2026) AIBOM runtime capture via OpenTelemetry: practice: runtime AIBOM generation using OpenTelemetry instrumentation
- [x] Mitchell (2026) Production incidents linked to AI systems: documented AI production incident failure modes and mitigations
- [x] Mitchell (2026) AI regulatory guidance delta check: new regulatory guidance and coverage gaps since prior review
- [x] Mitchell (2026) Five Eyes stance on AI risk and policy advice: Five Eyes AI risk posture and architectural security implications
- [x] Mitchell (2026) Open-weight model safeguard policy enforcement: safeguard and policy enforcement for open-weight models
- [x] Mitchell (2026) IT system legibility and measurement frameworks: legibility and measurement frameworks for IT systems; informs observability layer
- [x] Mitchell (2026) AI agent control-plane architecture in the enterprise: prior control-surface architecture used to qualify placement decisions
- [x] Mitchell (2026) AI agent identity and access management in the enterprise: prior identity and delegation architecture relevant to second-cycle identity findings
- [x] Mitchell (2026) Permission-safe Retrieval-Augmented Generation enterprise information architecture: prior retrieval and authorization architecture relevant to incident and observability conclusions
- [x] Mitchell (2026) Knowledge curation governance for regulated Artificial Intelligence: prior authoritative-source governance item that qualifies incident and evidence recommendations
- [x] Open Worldwide Application Security Project AIBOM: authoritative public definition surface for the AI bill-of-materials concept
- [x] CycloneDX AI/ML-BOM: standards-aligned definition and scope for AI and machine-learning bill-of-materials work
- [x] Backstage Software Catalog: authoritative catalog-coverage and ownership surface used in the legibility definition
- [x] ServiceNow CMDB Health: authoritative configuration-data quality surface used in the legibility definition
- [x] Dynatrace Smartscape: authoritative runtime-topology surface used in the legibility definition
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-09 | 4901d2e | Initial completion |