Updating the enterprise Artificial Intelligence ecosystem capability reference…

Updating the enterprise Artificial Intelligence ecosystem capability reference architecture using second-cycle 2026-05 completed items

2026-05-08 · ai-architecture governance-policy security-risk benchmarks-eval agentic-ai · synthesis medium · source → · wiki →
key claims
  1. The existing five-layer architecture remains usable, but the provenance plane now has to cover legibility, meaning catalog, configuration, topology, and drift visibility, and runtime evidence in addition to design-time lineage if it is to explain what the system actually became in operationBackstage (n.d.)ServiceNow (n.d.)Dynatrace (n.d.)Mitchell (2026)Mitchell (2026)
  2. Delivery and platform engineering now need an explicit declared AIBOM capability, where AIBOM is an AI bill-of-materials artifact for models, data, configuration, prompts, tools, and related execution context, because both managed platforms and code-native orchestration frameworks require deliberate extraction paths before approved design can be compared with later runtime behaviorOwaspaibom (n.d.)CycloneDX (n.d.)Mitchell (2026)Mitchell (2026)Mitchell (2026)
  3. Runtime-observed evidence has to become a first-class architecture service, because prompts, retrieved context, tool order, authority handoffs, and missing observability are governance-relevant facts that only appear after execution beginsMitchell (2026)Mitchell (2026)Mitchell (2026)
  4. The orchestration layer now needs explicit delegation-aware identity and tool-action components, because portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receiptMitchell (2026)Mitchell (2026)Mitchell (2026)
  5. Operating assurance now needs named architecture components for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weakMitchell (2026)Mitchell (2026)Mitchell (2026)
  6. Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architectureMitchell (2026)Mitchell (2026)Mitchell (2026)
  7. Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layersMitchell (2026)Github (n.d.)Github (n.d.)
  8. Central ownership should stay with provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds tied to local business contextMitchell (2026)Mitchell (2026)Mitchell (2026)

Research Question

How should the enterprise Artificial Intelligence (AI) ecosystem capability reference architecture (as expressed in 2026-04-22-enterprise-ai-capability-model, 2026-05-05-enterprise-ai-capability-stack, and 2026-05-06-ai-capability-reference-architecture-security-supply-chain-update) be revised and extended to incorporate findings from the second-cycle 2026-05 completed items, covering AI Bill of Materials (AIBOM) declared construction practices, effectiveness and risk-mitigation limits, multi-agent identity attribution, platform observability controls, European Union (EU) AI Act regulatory intersection, and OpenTelemetry-based runtime capture; AI production incidents; regulatory guidance updates; Five Eyes AI risk posture; open-weight model safeguard policies; and Information Technology (IT) legibility and measurement frameworks, that were not incorporated in the first-cycle update?

Findings

(Populated from §6 Synthesis above.)

Executive Summary

The enterprise AI reference architecture should remain a five-layer model, but the shared provenance and governance planes now need explicit second-cycle services for legibility, meaning catalog, configuration, topology, and drift visibility, runtime evidence, operational assurance, and regulator-facing governance.

Second-cycle evidence shows that design-time inventory alone is insufficient, because the architecture also has to preserve runtime divergence, delegation receipts, incident-triggering conditions, and estate legibility across catalogs, topology, and structural-drift evidence.

Regulatory and Five Eyes updates also move operating-model work into the architecture itself, because board literacy, supplier-risk management, secure logging, rollback authority, and evidence assembly are now observable control surfaces instead of background management assumptions.

The practical revision is to widen the provenance plane into a provenance-and-legibility plane, widen the governance plane into a governance, evaluation, and operational-assurance plane, and keep shared semantics and retained evidence under explicit central ownership.

Key Findings

  1. The existing five-layer architecture remains usable, but the provenance plane now has to cover legibility, meaning catalog, configuration, topology, and drift visibility, and runtime evidence in addition to design-time lineage if it is to explain what the system actually became in operation.
  2. Delivery and platform engineering now need an explicit declared AIBOM capability, where AIBOM is an AI bill-of-materials artifact for models, data, configuration, prompts, tools, and related execution context, because both managed platforms and code-native orchestration frameworks require deliberate extraction paths before approved design can be compared with later runtime behavior.
  3. Runtime-observed evidence has to become a first-class architecture service, because prompts, retrieved context, tool order, authority handoffs, and missing observability are governance-relevant facts that only appear after execution begins.
  4. The orchestration layer now needs explicit delegation-aware identity and tool-action components, because portable attribution breaks when systems change credential type, cross trust boundaries, or execute under shared runtime identities without a surviving delegation receipt.
  5. Operating assurance now needs named architecture components for authoritative-source binding, fairness validation, rollback authority, and external challenge escalation, because public harm repeatedly surfaced where systems looked trustworthy but source governance, deployment constraints, or oversight loops were too weak.
  6. Regulatory and Five Eyes updates mean that board literacy, supplier and concentration-risk governance, secure-by-design logging, input control, and regulator-facing evidence assembly must now be named operating-model capabilities rather than assumed management practices outside the architecture.
  7. Open-weight safeguard models fit best as shared governance and evaluation services that delivery gates or higher-risk runtime checkpoints can invoke, because they offer organization-specific policy judgment while remaining too infrastructure-heavy for standard hosted-runner paths and too narrow to replace deterministic or human review layers.
  8. Central ownership should stay with provenance schema, telemetry normalization, retained evidence, policy semantics, regulatory reporting, and evaluation standards, while domain teams keep responsibility for authoritative content, workflow composition, and risk-tuned thresholds tied to local business context.

Assumptions

Analysis

The second-cycle evidence makes the architecture more operational, not more abstract, because the missing pieces are evidence-handling and control-surface capabilities that only appear when systems are released, observed, challenged, and regulated in practice.

The main trade-off is between central coherence and layer-local enforcement. Centralizing everything would hide where trust decisions really occur, while leaving every team to invent its own evidence model would destroy comparability, auditability, and incident reconstruction.

The best-supported resolution is to keep control semantics and retained evidence centralized while leaving enforcement near the layer that owns the relevant trust decision, retrieval permissions in data and knowledge, semantic safeguards near inference, tool and delegation controls in orchestration, and release-time signing, extraction, and evaluation in delivery and platform engineering.

A major alternative would be to keep every second-cycle addition inside the existing layers without widening either cross-cutting plane, or to replace the baseline stack with a flatter capability mesh, but the evidence weighs against both moves because the new requirements cluster around shared evidence, legibility, and retained-governance services that cut across every layer rather than belonging cleanly to one local component family.

Plausible alternative placements exist for open-weight safeguard capability, especially model-layer controls and pre-deployment review gates, but the evidence supports locating the canonical policy service in the governance and evaluation plane so that delivery gates and runtime checkpoints can invoke one shared policy logic rather than duplicate policy semantics in each layer.

Revised component list:

  1. Data and knowledge layer: authoritative-source registry, permission-safe retrieval, retrieval-snapshot metadata, and upstream provider-disclosure links.
  2. Model and inference layer: approved model and provider registry, semantic safeguards, policy-conditioned classification hooks, and model-facing runtime signal capture.
  3. Orchestration and execution layer: tool allowlists, delegation-receipt capture, action checkpoints, recursion and stop authority, and per-run authority context.
  4. Delivery and platform-engineering layer: declared AIBOM builder, signed artifact lineage, promotion-time evaluation gates, telemetry collector pipeline, and approved-versus-observed divergence classifier.
  5. Operating-model layer: board and executive literacy, supplier and concentration-risk governance, fairness review, incident response, rollback authority, and regulator-facing evidence assembly.
  6. Provenance-and-legibility plane: declared and observed supply-chain records, ownership-bearing catalogs, architecture blueprints, runtime topology, and structural-drift evidence.
  7. Governance, evaluation, and operational-assurance plane: policy semantics, evaluation standards, exceptions, explainability artifacts, incident records, and optional second-stage safeguard review.

Ownership recommendations:

  1. Central platform, security, and risk functions should own provenance schema, telemetry normalization, retained evidence, evaluation standards, and regulatory reporting because those assets lose value when fragmented by domain.
  2. Domain teams should own authoritative content, workflow composition, business-specific thresholds, and local exception context because those choices depend on business meaning and risk appetite that the central platform cannot infer.
  3. Shared services should expose signed interfaces and evidence contracts rather than monolithic approval queues, because second-cycle evidence favors strong shared semantics with layer-local enforcement and action-specific stop authority.

Risks, Gaps, and Uncertainties

Open Questions


sources


cites
cites Enterprise AI capability model for use-case maturity decisions
cites Knowledge curation governance as an enterprise AI capability in regulated financial institutions
cites What control-plane architecture is required to manage Artificial Intelligence (AI) agents and low-code systems as distributed, semi-autonomous actors within enterprise environments?
cites What identity and access management model is required for Artificial Intelligence (AI) agents and low-code artefacts operating within enterprise systems?
cites Permission-safe Retrieval-Augmented Generation (RAG) in enterprise information architectures: technical constraints, architectural options, and failure modes at scale
cites Integrating 2026-05 security and supply chain findings into the enterprise Artificial Intelligence capability reference architecture
cites How do you construct a declared design-time Artificial Intelligence Bill of Materials (AIBOM) for a real tool-using, stateful Artificial Intelligence (AI) workload? A worked example using Amazon Web Services (AWS) Bedrock Agents and LangGraph
cites What security and governance risks can a declared and runtime-observed inventory of models, prompts, retrieval sources, tools, memory, and delegation artifacts realistically mitigate for tool-using, stateful Artificial Intelligence (AI) workloads, and where does it create false assurance?
cites How do OAuth 2.0, OpenID Connect, and SPIFFE token propagation work in real multi-agent pipelines, and where does end-to-end attribution break in practice?
cites What introspection, export, and control surfaces actually exist across production agentic Artificial Intelligence (AI) platforms: a comparative analysis of Amazon Web Services (AWS) Bedrock Agents, Microsoft 365 Copilot, Salesforce Agentforce, and ServiceNow Now Assist?
cites How does the European Union (EU) AI Act and related international AI governance regulation intersect with machine-readable AI component-inventory requirements for high-risk multi-step tool-using Artificial Intelligence (AI) systems?
cites How do you capture a runtime-observed Artificial Intelligence Bill of Materials (AIBOM) in practice using OpenTelemetry tracing and platform-native observability tools?
cites Production incidents linked to Artificial Intelligence systems
cites Artificial Intelligence (AI) regulatory guidance delta check: new advice, policy, and missed coverage since prior global financial-services review
cites Five Eyes stance on Artificial Intelligence risk and policy advice
cites How should human-in-the-loop (HITL) design be adapted when AI review volume makes human reviewers a bottleneck or causes rubber-stamping?
cites How do open-weight policy enforcement reasoning models, exemplified by OpenAI's gpt-oss-safeguard, classify text against customizable policies, and what are their deployment trade-offs compared to rule-based and closed Application Programming Interface (API) guardrail approaches?
cites What measurement systems and frameworks exist for quantifying Information Technology system legibility, the ability to reason about, understand, and comprehensively characterise a runtime ecosystem of interconnected applications, services, and systems, who is actively defining and applying them, and how?
related (frontmatter)
related Enterprise capability stack for sustainable multi-provider Artificial Intelligence
version history
versiondatecommitsummary
1.02026-05-094901d2eInitial completion

Connected items

Loading…

View full knowledge graph →