How Do Enterprise AI Maturity Frameworks Map onto the LLM Consumption Ladder?
How do theoretical frameworks of enterprise Artificial Intelligence (AI) / generative AI maturity map onto the observed, practice-driven progression of large language model (LLM) consumption strategie…
Secure Runtime Evolution for AI Coding Agents
What is the logical progression in AI (Artificial Intelligence) coding-agent runtime design from local process/Operating System (OS) sandboxes, through shared Continuous Integration (CI)/cloud develop…
SRE: establishing SLOs as contractual capability boundaries
How do Site Reliability Engineering (SRE) practices establish what a system can safely do, expressed as a contractual boundary rather than an observed average: specifically, how are Service Level Obje…
ITIL capacity management
What does IT Infrastructure Library (ITIL) capacity management specify as the measurement practice for establishing a platform capability baseline, and where does it rely on assertion rather than tele…
Flexibility vs. Predictability
In a production pipeline with uncontrolled inputs, how does the trade-off between the flexibility of an agentic system and the predictability of a deterministic execution model affect the auditability…
Structural Stability vs. Predictive Fragility
Using dynamical systems theory, how does the fragility of a purely predictive model under input noise or system drift differ from the local qualitative stability of a model whose governing equations p…
Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic…
What are the operational failure modes of an epistemic framework that prioritises instrumentalism, treating predictive performance as the primary criterion, over explanatory reach when applied to comp…
What Are We Losing and Gaining by Inserting Autonomous Tool-Using Artificial…synthesis
What are we concretely losing and gaining, across the dimensions of capability, reliability, auditability, explainability, and organisational risk, by inserting autonomous tool-using Large Language Mo…
Are Multi-Step Large Language Model-Based Systems Inherently Less Explainable…synthesis
Are multi-step Large Language Model (LLM)-based systems inherently less explainable than equivalently scoped deterministic software systems, or does production-scale distributed-system complexity make…
Longitudinal persistence rates after gap closure for low-code applications,…
What longitudinal evidence exists on persistence rates after the original gap is closed for low-code applications, bots, and agents in live enterprise estates?
Matched denominator for comparing post-pipeline release-based failures with…
What common denominator enables direct matched comparison between post-pipeline release-based failure rates and production live-runtime incident rates for the same production workflow?
Microsoft Copilot Studio
What is the complete set of features, functions, and capabilities offered by Microsoft Copilot Studio, and how do those capabilities support enterprise-grade Artificial Intelligence (AI) agent develop…
Microsoft Foundry (formerly Azure Artificial Intelligence (AI) Foundry)
What is the complete set of features, functions, and capabilities offered by Microsoft Foundry, and how do those capabilities support the full Artificial Intelligence (AI) development lifecycle, from…
Amazon Bedrock AgentCore and related suite
What is the complete set of features, functions, and capabilities offered by Amazon Bedrock AgentCore and its related suite, including AgentCore Gateway, AgentCore Memory, AgentCore Identity, and the…
External Dependency Surface Taxonomy for Production LLM Agents
What is the complete taxonomy of external dependencies for a production Large Language Model (LLM)-based agent, how does each dependency class fail, what is the blast radius of each failure class, and…
When Retrieval-Augmented Generation source documents change after agent build…
When the source documents indexed in a Retrieval-Augmented Generation (RAG) pipeline change after an agent has been built and tested, what failure modes and behavioral regressions can result in produc…
Data Governance Standards and Regulations Applied to Artificial Intelligence…
How do established data governance standards, including International Organization for Standardization and International Electrotechnical Commission (ISO/IEC) 38505, DAMA-DMBOK (Data Management Body o…
Build vs improve tradeoff
Given constrained engineering capacity, how should organisations allocate effort between (1) building features within an existing system and (2) improving the system itself (tooling, process, architec…
What tiered human oversight models maintain meaningful human-in-the-loop (HITL)…
Under high-volume deployment of multi-step Artificial Intelligence (AI) systems, what factors cause human-in-the-loop (HITL) oversight to degrade into rubber-stamping, meaning approval without genuine…
Production incidents linked to Artificial Intelligence systems
What documented production incidents over the last five years were caused or materially contributed to by Artificial Intelligence (AI) systems, and what recurring failure modes and mitigations were id…
What are the capabilities, architectural assumptions, and practical deployment…
What are the capabilities, underlying architectural assumptions, and practical deployment constraints of Loki as an MIT-licensed automated fact-checking tool optimised for journalists and content mode…
How do open-weight policy enforcement reasoning models, exemplified by OpenAI's…
How do open-weight, meaning released-weight and self-hostable, policy enforcement reasoning models, exemplified by OpenAI's gpt-oss-safeguard, classify text against strict, customizable policies, and…
Explainable Artificial Intelligence (XAI)
What is the current state of Explainable Artificial Intelligence (XAI) research, who leads it and what are the primary techniques, and how does XAI intersect with regulatory obligations, audit require…
Large Language Model (LLM)-as-judge as pipeline validation checkpoints
Which organisations, projects, and frameworks are defining and operationalising Large Language Model (LLM)-as-judge evaluation, the use of one model to assess another model's outputs, as automated val…
Alternative Continuous Integration and Continuous Delivery pipeline platforms…
What alternative Continuous Integration and Continuous Delivery (CI/CD) pipeline platforms, specifically Harness, Amazon Web Services (AWS) CodeBuild and CodeDeploy, and Jenkins, can serve as the gove…
Universal Entity Lifecycle Governance Framework (UELGF)
How should the UELGF specify the runtime feedback loop, covering signal taxonomy, signal aggregation and evaluation mechanism, automated response taxonomy proportionate to signal severity, re-evaluati…
What is the precise technical distinction between code generation and other…
What is the precise technical distinction between code generation and other Large Language Model (LLM)-generated outputs in terms of external verifiability, specifically, that code operates in a forma…
When and how should human intervention be incorporated into Artificial…
When and how should human intervention be incorporated into AI-driven and automated workflows, specifically, what trigger conditions, intervention thresholds, escalation procedures, response time expe…
How should AI and low-code governance integrate with existing software…
How should Artificial Intelligence (AI) and low-code governance integrate with existing software development and platform engineering practices, specifically, how should governance controls be integra…
What observability and telemetry model is required to govern Artificial…
What observability and telemetry model is required to govern AI and low-code systems at scale, specifically, what must be logged, at what frequency, and at what level of granularity, including prompt…
What lifecycle management model is required for Artificial Intelligence (AI)…
What comprehensive lifecycle management model is required for AI models, prompts, and low-code applications, covering versioning strategies, deployment controls, rollback mechanisms, ownership trackin…
What are the primary failure modes in enterprise Artificial Intelligence (AI)…
What are the primary failure modes in enterprise Artificial Intelligence (AI) and low-code deployments, including data leakage, conflicting automations, unintended actions by AI agents, and loss of au…
How should decision rights, accountability, and liability be structured for…
How should decision rights, accountability, and liability be structured for AI systems and low-code applications in enterprise environments, specifically, who should be empowered to approve new use ca…
Deployment pipeline as the only enforceable control gate for citizen-developed…
In an environment where citizen development tooling is already licensed and accessible to non-technical staff, and where the distinction between personal productivity and production automation has col…
Dependency ordering of foundational conditions for safe agentic Artificial…
The foundational conditions for safe agentic AI deployment in a regulated financial institution are not independent, they form a dependency graph in which policy coherence is a prerequisite for inform…
Enterprise AI use-case routing frameworks
What decision frameworks do enterprises use to route Artificial Intelligence (AI) use cases to the appropriate platform, implementation pattern, and risk tier, distinguishing low-code business-led, pr…
Enterprise AI platform operating models
What organisational structures do enterprises use to operate multiple Artificial Intelligence (AI) platforms simultaneously, and what trade-offs emerge between (a) a single unified AI platform team, (…
Automated governance assurance and change control verification patterns for…
What technical patterns exist for automating governance assurance and change control verification in Artificial Intelligence (AI)-assisted delivery pipelines, specifically audit evidence generation, p…
Enterprise AI capability model for use-case maturity decisions
What enterprise-wide Artificial Intelligence (AI) capability model best supports deciding whether a candidate AI use case requires net-new foundational capabilities or can reuse capabilities already b…
Claude Code npm Source Map Leak
How did the March 2026 accidental leak of Anthropic's Claude Code source code via an npm (Node Package Manager) package occur, and what processes and protections can organisations adopt to prevent sim…
Backpressure Infrastructure and the Theory of Constraints
What is backpressure infrastructure, specifically as it pertains to the Theory of Constraints (TOC), and what does academic research and real-world white papers say about its practical application?
Environment setup consistency
Given the two primary agent entry points, (A) assigning a GitHub issue to the Copilot coding agent and (B) using the Claude iOS `code` feature, what environment does each agent start in, and what cont…
Applied context engineering
What practical patterns, workflow best practices, and agent development guidelines emerge from synthesising the `muratcankoylan/Agent-Skills-for-Context-Engineering` skill library with the context eng…
Stateless-agent assumption failure
When an agentic workflow spans multiple session boundaries — each session starting with a fresh context window and no memory of prior runs — what are the mechanisms by which external state becomes orp…
Hosting options for the Research repo
What is the best free or very-low-cost hosting option for this research repository that supports full-text search, and optionally vector/graph database capabilities, without requiring SEO, custom DNS,…
Failure mode taxonomy
The five-layer failure mode taxonomy established in `2026-03-10-ai-concept-classification-taxonomy.md` (Q5) provides a structurally sound classification, but leaves three empirical gaps unanswered: (1…
Self-hosted MCP server options
What is the minimum viable self-hosted deployment of `mcp_server.py` (or a write-only HTTP wrapper) that: (a) is reachable from the public internet, (b) has zero or near-zero ongoing cost, (c) require…
LanceDB index rebuild speed from git
Can the LanceDB index be rebuilt from the `.md` files in the repo on startup fast enough to enable stateless (per-request) deployment? Measure rebuild time at: current corpus size, 100 files, 500 file…
RBNZ AI Supervisory Expectations
What are the Reserve Bank of New Zealand's specific supervisory expectations for AI use by regulated entities, and how do these align with or diverge from the expectations of comparator regulators (AP…
Machine Learning (ML) technique taxonomy and selection criteria for analytics…
What is the complete, structured landscape of machine learning techniques and algorithms that an advanced analytics department should know, use, and actively pursue — covering foundational concepts, w…