Macro-level hallucination risk in schema-free GraphRAG clustering
How does the noisy baseline produced by unconstrained entity extraction corrupt the hierarchical summaries generated by standard Graph Retrieval-Augmented Generation (GraphRAG) community-detection pip…
Context collision and relational blindness in flat-vector RAG
Given that classical flat-vector Retrieval-Augmented Generation (RAG) acts as an external access mechanism rather than a persistent internal memory state, how do contradictory semantic overlaps in top…
What constitutes cohesive and coherent organisational governance for aligned,…
What constitutes good cohesive and coherent organisational governance, meaning the specific configurations, principles, mechanisms, and performance thresholds that reliably produce aligned, high-veloc…
Governance and operating models for safe-to-fail experimentation in regulated…
In highly regulated industries such as financial services, healthcare, and pharmaceuticals, how do organisations design governance structures, team models, and operating practices that enable safe-to-…
Evaluation frameworks for agentic memory quality, relevance, and retrieval…
What benchmark suite and metric design best measures the quality, relevance, retrieval accuracy, freshness, and governance correctness of agentic memory systems across heterogeneous tasks?
TBox-driven vs ABox-emergent ontology approaches in GraphRAG systems
To what extent do TBox (Terminological Box)-driven (predefined upper- and mid-level) ontologies outperform, underperform, or complement ABox (Assertion Box)-emergent (bottom-up, data-driven) approache…
How should the balance between standardized and customized internal tooling…synthesis
How should the balance between standardized and customized internal tooling shift across industries, organisation sizes, maturity levels, and Artificial Intelligence (AI) agent adoption patterns, and…
At what scale or under what operating conditions do the aggregate costs of…
At what scale or under what operating conditions do the aggregate costs of fragmented local tooling exceed the productivity gains from customization, and which metrics let organisations detect that cr…
AI productivity, quality, and governance open questions
What empirical evidence can distinguish sustainable Artificial Intelligence (AI)-enabled software delivery gains from short-lived throughput effects and hidden quality or governance costs in productio…
Capability claim vs. production telemetry
When a team's capability claim conflicts with production telemetry, what arbitration mechanism produces a reliable baseline, and is there empirical evidence on which approach (telemetry override, stru…
How have software-development commit trends shifted across repository creation,…
What do high-quality longitudinal studies (2019–2026) show about directional shifts and current baseline ranges for repository creation rate, Lines of Code (LOC) velocity, rework share, project abando…
Q6: Leading indicators of instability in split-authority flow systems
Which metrics best predict unsafe queue growth, rising delivery risk, or hidden demand accumulation in a split-authority delivery system, where "split-authority" means a context in which authority is…
Joint Embedding Predictive Architecture (JEPA) shift
Is the shift from text-token prediction to Joint Embedding Predictive Architecture (JEPA)-style video outcome prediction the same class of problem as the shift from video prediction to physically grou…
LLM reasoning in mathematics and programming tasks
To what extent is the claim true that mathematics and programming are especially strong use cases for Large Language Models (LLMs) because both rely on formal symbolic languages that may align with mo…
Similarity algorithms and growth policy for a file-based controlled theme…
Which similarity algorithms are appropriate for detecting near-synonym themes in a controlled vocabulary of 20–40 slug-based labels, and what growth policy prevents both vocabulary explosion and colla…
Theory and mechanisms of prompt and program optimization in Language Models
What theory best explains why prompt and program optimization methods can outperform baseline prompting and Reinforcement Learning (RL) in Language Model (LM) pipelines, and how do the methods in the…
Contract theory formulation and statistical criteria for contracts
What is contract theory, how is a contract-theory model formally formulated, what is meant by a statistical contract, meaning a contract or protocol whose payoffs depend on statistical evidence, and w…
How should financial Retrieval-Augmented Generation (RAG) systems filter…
What pre-retrieval architecture and governance controls most reliably remove low-information content, meaning boilerplate, repeated passages, wrapper text, and other low-signal document fragments, and…
Adversarial Input Propagation Through Multi-Step Tool-Using LLM Systems
How do adversarial inputs or unexpected environmental shifts propagate error through a multi-step tool-using Large Language Model (LLM) system's verification and strategy-selection phases when the und…
Agentic Tool-Feedback Loops and Explanatory Reach
When a Large Language Model (LLM) is wrapped in an agentic loop, meaning a repeated perception, strategy-selection, tool-action, and verification cycle, does the outer loop introduce true explanatory…
In-Context Learning and Chain-of-Thought Prompting
What are the empirical boundaries of in-context learning and chain-of-thought prompting when they are used to push a purely predictive statistical architecture toward intervention questions and altern…
The Stochastic Parrot Under Pressure
How does the Stochastic Parrot hypothesis, the claim that Large Language Models (LLMs) reproduce linguistic form more readily than grounded structural understanding, manifest when an LLM is presented…
Large Language Models as Statistical Optimisers
To what extent do Large Language Models (LLMs) optimise strictly for linguistic form and statistical token distribution rather than constructing internal, invariant causal models of reality?
The Duhem-Quine Thesis and Underdetermination
How can the phenomenon of multiple distinct functions perfectly interpolating identical data points be formalised through the lens of the Duhem-Quine thesis, underdetermination of theory by data, and…
Empirical Risk Minimisation's Causal Blindness
How does the framework of Empirical Risk Minimisation (ERM) mathematically guarantee predictive accuracy within a known data distribution while remaining blind to the stable cause-and-effect relations…
Failure Modes of Instrumentalist Epistemology When Applied to Complex Dynamic…
What are the operational failure modes of an epistemic framework that prioritises instrumentalism, treating predictive performance as the primary criterion, over explanatory reach when applied to comp…
David Deutsch's Hard-to-Vary Criterion
Using David Deutsch's hard-to-vary criterion, meaning an explanation whose details cannot be changed without losing explanatory force, what formal criteria can measure the internal logical constraints…
Formalising Popper's Falsifiability as a Mathematical Criterion for…
How can Karl Popper's criterion of demarcation and falsifiability be mathematically formalised to distinguish between a model that explains a physical mechanism and one that merely interpolates observ…
What Are We Losing and Gaining by Inserting Autonomous Tool-Using Artificial…synthesis
What are we concretely losing and gaining, across the dimensions of capability, reliability, auditability, explainability, and organisational risk, by inserting autonomous tool-using Large Language Mo…
Are Multi-Step Large Language Model-Based Systems Inherently Less Explainable…synthesis
Are multi-step Large Language Model (LLM)-based systems inherently less explainable than equivalently scoped deterministic software systems, or does production-scale distributed-system complexity make…
Datasets for measuring conversion from demand for local workaround tools to…
What public or internal datasets can validly measure the rate at which demand for local workaround tools, such as local apps, flows, lists, or spreadsheets, is converted into formal central Informatio…
Matched denominator for comparing post-pipeline release-based failures with…
What common denominator enables direct matched comparison between post-pipeline release-based failure rates and production live-runtime incident rates for the same production workflow?
Policy Quality Degradation and Cross-Institution Blind Spots When New Policy…
What policy-quality degradation and systemic blind-spot risks emerge when organisations draft new policy versions from Large Language Model (LLM) interpretations of previous policy versions?
LLM Response Style and Confidence Signalling
How do Large Language Model (LLM) response style and self-reported confidence change how accurately users judge uncertainty and downstream risk when interpreting ambiguous policy and compliance requir…
LLM Training Prior Contamination in Compliance Interpretation
What failure modes emerge when Large Language Models (LLMs) combine generic public legal knowledge with proprietary organisational policy in compliance interpretation tasks?
Cognitive Closure Under Ambiguity and Confirmation Bias
How do pressures to reach a quick, definite answer under ambiguity and iterative prompt refinement influence acceptance of flawed Large Language Model (LLM) policy interpretations?
De Facto Policy Drift From Repeated Unverified LLM Interpretations
How quickly do repeated unverified Large Language Model (LLM) interpretations create de facto policy norms that diverge from executive intent and board-level risk appetite?
Adversarial prompting risks in policy assistants
How vulnerable are corporate compliance Large Language Models (LLMs) to adversarial prompting that reframes restrictive policy as permissive guidance, and which controls detect or contain deliberate m…
Microsoft Foundry (formerly Azure Artificial Intelligence (AI) Foundry)
What is the complete set of features, functions, and capabilities offered by Microsoft Foundry, and how do those capabilities support the full Artificial Intelligence (AI) development lifecycle, from…
Amazon Web Services (AWS) Bedrock platform capabilities
What is the complete set of features, functions, and capabilities offered by Amazon Web Services (AWS) Bedrock, including its model access, agent building, knowledge bases, guardrails, evaluation, and…
Variance Control Comparison Across Delivery Modes
What is the empirical failure-rate distribution of Artificial Intelligence (AI)-assisted code that has passed a standard software delivery pipeline compared with AI-agent-executed business processes a…
Information Technology (IT) throughput capacity as a constraint on unmet…
What is the empirical relationship between Information Technology (IT) throughput capacity and the rate at which unmet operational capability needs accumulate across comparable organisations, and what…
External Dependency Surface Taxonomy for Production LLM Agents
What is the complete taxonomy of external dependencies for a production Large Language Model (LLM)-based agent, how does each dependency class fail, what is the blast radius of each failure class, and…
Temporary Automation Demand Persistence and Core Capability Investment…
What evidence exists that temporary automation workarounds displace investment in core software delivery, and what is the observed persistence rate of those workarounds after the underlying systems ca…
Agent Operational Cost vs Gap Closure Cost
What is the fully loaded operational cost of a production Artificial Intelligence (AI) agent used as a workaround for a missing system capability, relative to the cost of closing the underlying system…
Endsley Model of Situational Awareness deep dive
What is the Endsley Model of situational awareness, meaning the perception of relevant elements, comprehension of their meaning, and projection of their near-future status, how are its three levels de…
What is Anthropic's '4D' framework for Artificial Intelligence (AI) fluency,…
What is Anthropic's "4D" framework for Artificial Intelligence (AI) fluency, what do each of the four Ds, Delegation, Description, Discernment, and Diligence, mean in practice, and how does this frame…
Architectural patterns for reliable organizational process identification,…
What integrated architectural configuration of retrieval, reconciliation, constraint enforcement, memory, validation, escalation, and governance mechanisms most reliably enables visual workflow toolin…
When Retrieval-Augmented Generation source documents change after agent build…
When the source documents indexed in a Retrieval-Augmented Generation (RAG) pipeline change after an agent has been built and tested, what failure modes and behavioral regressions can result in produc…
Hardware load and Large Language Model (LLM) inference performance
How does hardware resource load, Central Processing Unit (CPU), Graphics Processing Unit (GPU), and memory pressure, affect Large Language Model (LLM) inference performance, specifically latency, thro…
Practical Limits of Large Language Model (LLM) Determinism
What are the practical limits of making LLM (Large Language Model)-based decisions or policy enforcement deterministic, even with temperature=0, fixed seeds, and constrained prompts?
Updating the enterprise Artificial Intelligence ecosystem capability reference…synthesis
How should the enterprise Artificial Intelligence (AI) ecosystem capability reference architecture (as expressed in `2026-04-22-enterprise-ai-capability-model`, `2026-05-05-enterprise-ai-capability-st…
Integrating 2026-05 security and supply chain findings into the enterprise…synthesis
How should the enterprise Artificial Intelligence (AI) ecosystem capability reference architecture (as expressed in `2026-04-22-enterprise-ai-capability-model` and the `2026-05-05-enterprise-ai-capabi…
What is the architecture and practical applicability of OpenFactCheck as an…
What is the architecture, evaluation methodology, and practical applicability of OpenFactCheck as an automated, modular, claim-level fact-checking pipeline for Artificial Intelligence (AI)-generated c…
What are the capabilities, architectural assumptions, and practical deployment…
What are the capabilities, underlying architectural assumptions, and practical deployment constraints of Loki as an MIT-licensed automated fact-checking tool optimised for journalists and content mode…
How do open-weight policy enforcement reasoning models, exemplified by OpenAI's…
How do open-weight, meaning released-weight and self-hostable, policy enforcement reasoning models, exemplified by OpenAI's gpt-oss-safeguard, classify text against strict, customizable policies, and…
How does Factual precision Scoring (FActScore) operationalise atomic-level…
How does FActScore (Factual precision Scoring), developed at the University of Washington, operationalise the concept of atomic factual claim decomposition and precision scoring for Large Language Mod…
How can findings from OpenFactCheck, Loki, FActScore, gpt-oss-safeguard, and…synthesis
How can the findings from research into OpenFactCheck, Loki, FActScore, gpt-oss-safeguard, and Barnum statement identification techniques be synthesised into concrete, actionable improvements to the a…
What are Barnum statements (Forer Effect statements), how do they manifest in…
What are Barnum statements (also known as Forer Effect statements) as a class of vague, universally applicable assertions, how do they manifest specifically in Artificial Intelligence (AI)-generated r…
What measurement systems and frameworks exist for quantifying Information…
What measurement systems and frameworks exist for quantifying Information Technology (IT) system legibility, defined here as the ability to reason about, understand, and comprehensively characterise t…
What does the 2026 Harvard Business Review trendslop study and related…
What does the March 2026 Harvard Business Review (HBR) "trendslop" study reveal about positional bias, prompt-framing sensitivity, and context-insensitive bias in Artificial Intelligence (AI)-generate…
What systematic review methodologies and Artificial Intelligence (AI)-assisted…
What systematic review methodologies, Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA), Cochrane review, narrative synthesis, meta-ethnography, and realist synthesis, and wh…
How does STORM's perspective discovery step work, and what is the…
How does the STORM (Synthesis of Topic Outlines through Retrieval and Multi-perspective question generation) system's perspective discovery step generate diverse expert viewpoints before decomposing a…
What structured approaches and Artificial Intelligence (AI) agent workflow…
What structured approaches, from academic writing pedagogy, Artificial Intelligence (AI)-assisted writing tools, and agent workflow design, exist for converting synthesised research findings into poli…
What automated claim verification approaches against scientific literature…
What automated claim verification approaches against scientific literature, specifically arXiv preprints, are used in research synthesis systems, what search strategies maximise recall and precision f…
What adversarial review and red-teaming methods are most effective for…
What adversarial review and red-teaming methods, drawn from Artificial Intelligence (AI) safety research, debate-based evaluation, formal argumentation theory, and scientific peer review practice, are…
What does TerminalBench reveal about minimal toolsets and coding agent…
What does the TerminalBench benchmark reveal about the relationship between toolset minimalism and coding agent performance, and what design principles does it suggest for effective Artificial Intelli…
How do errors compound in Artificial Intelligence (AI)-agent-heavy codebases,…
How do errors ("boooos") compound in codebases developed with high volumes of AI agent-generated code, including how local patches cause global regressions, and what review and governance strategies c…
What criteria define tasks where Artificial Intelligence (AI) coding agents…
What empirically grounded criteria define the characteristics of software development tasks where Artificial Intelligence (AI) coding agents reliably add value, versus tasks where agent autonomy intro…
Artificial Intelligence coding harness quality benchmarks
What benchmarks, metrics, and evaluation methodologies are used to measure the quality of Artificial Intelligence (AI) coding harnesses, including Integrated Development Environment (IDE) plugins, age…
Test-Driven Development (TDD) and fast feedback loops in Artificial…
How does enforcing Test-Driven Development (TDD) with AI coding assistants, writing failing tests before asking the AI to implement, change the quality and stability of the AI output compared to "writ…
Software Engineering fundamentals and AI code generationsynthesis
Drawing on the planned seven-item research programme on Software Engineering (SE) fundamentals in Artificial Intelligence (AI)-augmented development, six completed primary items plus external anchors…
Grill-Me technique: iterative structured interviewing for human and Artificial…
How effectively does the "Grill Me" technique, relentless iterative structured interviewing of the human developer by the AI assistant to build a shared design concept before generating any code, redu…
Fundamentals-first versus specs-to-code
What empirical patterns emerge when comparing real-world software projects built with a strict fundamentals-first Artificial Intelligence (AI) workflow, structured alignment, modules with simple inter…
Deep modules in AI-augmented development
How much more effective is Artificial Intelligence (AI) at understanding, navigating, and correctly modifying a codebase composed of deep modules with simple interfaces versus one filled with many sha…
Artificial Intelligence code entropy and complexity
Does repeated Artificial Intelligence (AI) code generation without strong architectural guardrails demonstrably increase software entropy and complexity over time, as predicted by the entropy model de…
Human cognitive bias toward Artificial Intelligence (AI) correctness and…
To what extent do humans systematically over-trust AI-generated explanations, and what mechanisms, automation bias, RLHF-induced sycophancy in post-training, and the polysemantic nature of internal mo…
Large Language Model (LLM)-as-judge as pipeline validation checkpoints
Which organisations, projects, and frameworks are defining and operationalising Large Language Model (LLM)-as-judge evaluation, the use of one model to assess another model's outputs, as automated val…
How do academic and scientific publishing systems handle post-publication…
How do established academic and scientific publishing systems (journal publishers, preprint servers, living review platforms) handle post-publication corrections, amendments, retractions, and formal c…
TimesFM and the Landscape of Time-Series Foundation Models
What are the practical use cases for TimesFM (Google's pretrained time-series foundation model), who is doing comparable work, and how does the foundation-model paradigm extend to other structured dat…
Applied context engineering
What practical patterns, workflow best practices, and agent development guidelines emerge from synthesising the `muratcankoylan/Agent-Skills-for-Context-Engineering` skill library with the context eng…
Prompt injection threat landscape
What is the current state of the prompt injection threat in agentic artificial intelligence (AI) systems: who is exploiting it, who is defending against it, and what does the research community consid…
Research loop evaluation rubric
What structured rubric should be used to evaluate the outputs of this repository's research loop agent — and what does a minimal viable implementation of a Continuous Integration (CI)-integrated eval…
Agent evaluation framework
What evaluation framework allows systematic comparison of agent implementations across multiple repositories — identifying what problems each is solving, whether concepts are used idiomatically or in…
Adversarial agents with shared goals
What is the design pattern for a system of agents — human or AI — that share a common goal but deliberately occupy different competency domains and time horizons? How does "adversarial collaboration"…
Coverage gaps in automated research review skills, peer review patterns for…
What review methodology is required to reliably move from information gathering through to applied knowledge and wisdom — and which of those steps can be automated in a CI pipeline versus requiring hu…
Self-improving Artificial Intelligence (AI) agent evaluation loop architecture
What is the most principled architecture for a Self-Improving AI Agent Evaluation Loop — specifically, how should a nested inner/outer loop be designed so that a "Meta-Optimizer" rewrites system promp…
LLM Hallucinations — Types, Causes, and Current Mitigation Approaches
What are the established types, root causes, and current mitigation strategies for hallucinations in large language models, and what does the macroscopic (training-level) view leave unexplained that m…
Research output types
What are the possible output types from a research item, and how should each type be handled, stored, and acted upon?
Evaluating and improving autonomous research loop quality
How can the quality of research items produced by the `research-loop.yml` autonomous pipeline be systematically evaluated, and what changes to `research-prompt.md` and the loop's prompting strategy wo…
Research agenda curation
How should the research backlog be maintained and prioritised to ensure balanced coverage of important domains, detect over-concentration in one area (research drift), and surface high-value gaps — ra…
Artificial Intelligence (AI) risk-reduction deployments in financial services
Which organisations have developed AI strategies explicitly framed around risk reduction — operational risk, credit risk, fraud, compliance, model risk — and what governance structures, outcome metric…