Theme: llm-reasoning

52 items

← all themes
Macro-level hallucination risk in schema-free GraphRAG clustering
2026-08-20
How does the noisy baseline produced by unconstrained entity extraction corrupt the hierarchical summaries generated by standard Graph Retrieval-Augmented Generation (GraphRAG) community-detection pip…
Context collision and relational blindness in flat-vector RAG
2026-08-20
Given that classical flat-vector Retrieval-Augmented Generation (RAG) acts as an external access mechanism rather than a persistent internal memory state, how do contradictory semantic overlaps in top…
Hybrid memory integration
2026-07-20
How can Artificial Intelligence (AI) agents effectively synchronize structured semantic memory, meaning ontologies and knowledge graphs, with latent knowledge encoded in Large Language Model (LLM) wei…
Episodic-to-semantic memory consolidation in AI agents
2026-07-20
What techniques enable AI agents to reliably generalize from specific episodic experiences (interaction logs, task traces, observed events) to durable semantic memory entries (ontological facts, proce…
Autonomous knowledge curation and truth maintenance for agentic ontologies
2026-07-20
What mechanisms exist, or are under active research, to enable Artificial Intelligence (AI) agents to autonomously curate which extracted knowledge is worth retaining in a long-term ontology, detect a…
TBox-driven vs ABox-emergent ontology approaches in GraphRAG systems
2026-07-20
To what extent do TBox (Terminological Box)-driven (predefined upper- and mid-level) ontologies outperform, underperform, or complement ABox (Assertion Box)-emergent (bottom-up, data-driven) approache…
Ontology Completeness as a World Model for Large Language Model (LLM) Prediction
2026-05-25
To what extent can a sufficiently complete ontology function as a practical world model (in the sense described by Yann LeCun) for Large Language Models (LLMs) making predictive inferences, and which…
LLM reasoning in mathematics and programming tasks
2026-05-25
To what extent is the claim true that mathematics and programming are especially strong use cases for Large Language Models (LLMs) because both rely on formal symbolic languages that may align with mo…
Theory and mechanisms of prompt and program optimization in Language Models
2026-05-21
What theory best explains why prompt and program optimization methods can outperform baseline prompting and Reinforcement Learning (RL) in Language Model (LM) pipelines, and how do the methods in the…
Stochastic LLM Agent vs. Deterministic Coded System
2026-05-18
How do the failure modes of a stochastic multi-step Large Language Model (LLM) agent, meaning a tool-using system whose action path can vary across runs, differ fundamentally from the failure modes of…
Formal Generalisation Bounds for Tool-Using LLM Systems When Tools Return…
2026-05-18
What formal bounds can be stated for generalisation outside the training distribution in tool-using Large Language Model systems when their tools return non-deterministic outputs under unconstrained p…
Agentic Tool-Feedback Loops and Explanatory Reach
2026-05-18
When a Large Language Model (LLM) is wrapped in an agentic loop, meaning a repeated perception, strategy-selection, tool-action, and verification cycle, does the outer loop introduce true explanatory…
In-Context Learning and Chain-of-Thought Prompting
2026-05-18
What are the empirical boundaries of in-context learning and chain-of-thought prompting when they are used to push a purely predictive statistical architecture toward intervention questions and altern…
The Stochastic Parrot Under Pressure
2026-05-18
How does the Stochastic Parrot hypothesis, the claim that Large Language Models (LLMs) reproduce linguistic form more readily than grounded structural understanding, manifest when an LLM is presented…
Large Language Models as Statistical Optimisers
2026-05-18
To what extent do Large Language Models (LLMs) optimise strictly for linguistic form and statistical token distribution rather than constructing internal, invariant causal models of reality?
Pearl's Causal Hierarchysynthesis
2026-05-18
What are the formal information-theoretic boundaries that prevent a model trained exclusively on observational data (Level 1 on Pearl's Ladder of Causation) from ever executing or predicting the outco…
David Deutsch's Hard-to-Vary Criterion
2026-05-18
Using David Deutsch's hard-to-vary criterion, meaning an explanation whose details cannot be changed without losing explanatory force, what formal criteria can measure the internal logical constraints…
LLM-First Policy Clarification and Institutional Knowledge Atrophy
2026-05-17
How does shifting from peer policy clarification to Large Language Model (LLM)-first interaction affect institutional memory transfer, mentoring, and long-term policy expertise?
Policy Quality Degradation and Cross-Institution Blind Spots When New Policy…
2026-05-17
What policy-quality degradation and systemic blind-spot risks emerge when organisations draft new policy versions from Large Language Model (LLM) interpretations of previous policy versions?
LLM Training Prior Contamination in Compliance Interpretation
2026-05-17
What failure modes emerge when Large Language Models (LLMs) combine generic public legal knowledge with proprietary organisational policy in compliance interpretation tasks?
Adversarial prompting risks in policy assistants
2026-05-17
How vulnerable are corporate compliance Large Language Models (LLMs) to adversarial prompting that reframes restrictive policy as permissive guidance, and which controls detect or contain deliberate m…
AI-Assisted Policy Interpretation and Accountability Displacement
2026-05-17
How does integration of Large Language Models (LLMs) into policy-ambiguity resolution change liability allocation, escalation behaviour, and an organisation's ability to justify the resulting decision…
Layered reasoning stack interfaces
2026-05-17
What state abstraction boundaries and interface protocols are most effective for mapping Large Language Model (LLM) candidate outputs into Energy-Based Model (EBM) evaluation state spaces while preser…
Agent-to-Agent (A2A)-to-tool-calling unification
2026-05-13
To what extent does unifying specialised Agent-to-Agent (A2A) protocols into a standardised tool-calling interface affect orchestration overhead and reasoning accuracy in hierarchical multi-agent syst…
Security, Compliance, and Governance Risks of Using Generative AI (GenAI) Tools…
2026-05-10
What are the documented security, compliance, and governance risks of using Generative Artificial Intelligence (GenAI) tools such as Microsoft 365 (M365) Copilot for drafting memos, reports, and other…
Practical Limits of Large Language Model (LLM) Determinism
2026-05-09
What are the practical limits of making LLM (Large Language Model)-based decisions or policy enforcement deterministic, even with temperature=0, fixed seeds, and constrained prompts?
Hybrid Architecture Design
2026-05-09
How should hybrid architectures be designed so that probabilistic LLMs handle interpretation and insight generation while deterministic layers enforce final governance, compliance, and high-stakes dec…
Compliance Risks of Relying on Stochastic Large Language Model (LLM) Outputs…
2026-05-09
What evidence or guidance exists on the compliance risks of relying primarily on stochastic Large Language Model (LLM) outputs for governance, privacy, or regulatory decisions?
How do open-weight policy enforcement reasoning models, exemplified by OpenAI's…
2026-05-06
How do open-weight, meaning released-weight and self-hostable, policy enforcement reasoning models, exemplified by OpenAI's gpt-oss-safeguard, classify text against strict, customizable policies, and…
How does Factual precision Scoring (FActScore) operationalise atomic-level…
2026-05-06
How does FActScore (Factual precision Scoring), developed at the University of Washington, operationalise the concept of atomic factual claim decomposition and precision scoring for Large Language Mod…
How can findings from OpenFactCheck, Loki, FActScore, gpt-oss-safeguard, and…synthesis
2026-05-06
How can the findings from research into OpenFactCheck, Loki, FActScore, gpt-oss-safeguard, and Barnum statement identification techniques be synthesised into concrete, actionable improvements to the a…
What are Barnum statements (Forer Effect statements), how do they manifest in…
2026-05-06
What are Barnum statements (also known as Forer Effect statements) as a class of vague, universally applicable assertions, how do they manifest specifically in Artificial Intelligence (AI)-generated r…
How does STORM's perspective discovery step work, and what is the…
2026-05-02
How does the STORM (Synthesis of Topic Outlines through Retrieval and Multi-perspective question generation) system's perspective discovery step generate diverse expert viewpoints before decomposing a…
What structured approaches and Artificial Intelligence (AI) agent workflow…
2026-05-02
What structured approaches, from academic writing pedagogy, Artificial Intelligence (AI)-assisted writing tools, and agent workflow design, exist for converting synthesised research findings into poli…
What adversarial review and red-teaming methods are most effective for…
2026-05-02
What adversarial review and red-teaming methods, drawn from Artificial Intelligence (AI) safety research, debate-based evaluation, formal argumentation theory, and scientific peer review practice, are…
Grill-Me technique: iterative structured interviewing for human and Artificial…
2026-04-30
How effectively does the "Grill Me" technique, relentless iterative structured interviewing of the human developer by the AI assistant to build a shared design concept before generating any code, redu…
Is knowledge scaffolding an established concept within context engineering for…
2026-04-29
Is knowledge scaffolding an established concept within context engineering for Large Language Models (LLMs) and Artificial Intelligence (AI) agents, and if so, how is it defined, implemented, and dist…
What is Yann LeCun's complete argument against Large Language Models as a path…
2026-04-26
What is Yann LeCun's complete and precise argument against Large Language Models (LLMs) as a path to autonomous machine intelligence, meaning Artificial Intelligence (AI) that can reason, plan, and ac…
Harness-level selection and use of tools, agents, skills, prompts, and…
2026-04-20
When should teams choose tools, agent definition files, skills, prompts, instruction files, and AGENTS.md, and what verifiable best practices align with how major harnesses actually select and apply e…
Anthropic Claude Code leak
2026-04-02
What does the accidental March 2026 leak of Anthropic's Claude Code source code reveal about: (1) the codebase architecture, (2) how key engineering problems are solved, (3) the prompting and instruct…
Large Language Models as offensive security tools
2026-03-31
What is the current state of Large Language Model (LLM)-driven offensive security capability: can LLMs autonomously discover and exploit zero-day (0-day) vulnerabilities, what does the empirical evide…
The role of AGENTS.md in a repo using .github/copilot-instructions.md as the…
2026-03-29
`AGENTS.md` has emerged as the cross-tool convergence format for agent project instructions, supported by OpenAI Codex, GitHub Copilot, Claude Code, Cursor, Aider, Gemini Command Line Interface (CLI),…
Working memory architecture, prefrontal cortex contextual gating, and…
2026-03-22
How do human brains store, compress, retrieve, and dynamically layer multiple types of contextual knowledge — values, goals, rules, current state, and immediate task — when making decisions, and what…
More formal proof engineering
2026-03-22
What does Leanstral - an open-source agent for formal proof engineering - offer as a practical path to trustworthy, formally verified software built with Artificial Intelligence (AI) assistance, and h…
Reliable Software in the LLM Era
2026-03-16
What strategies and formal-methods tooling exist for maintaining software reliability in the Large Language Model (LLM) era, and what does the Quint formal specification language ecosystem - including…
AI concept classification taxonomy
2026-03-10
What is a coherent, internally consistent classification taxonomy for the core concepts in AI-assisted and agentic systems — covering prompt types, instruction types, prompt/content/intent engineering…
Context engineering: first principles of steering LLM output without control
2026-03-09
What are the first principles of context engineering — and what novel approaches emerge when it is understood as two distinct but coupled mechanisms: (1) making the next predicted token more likely to…
Emergent Patterns in Software Engineering Prompts and SDLC Guidance
2026-03-08
What are the current and emergent best practices for crafting AI agent prompts and tooling guidance tailored to each phase of the Software Development Life Cycle (SDLC) — covering discovery, requireme…
Pre-Training Origins of Hallucination-Associated Neurons — Implications for LLM…
2026-03-06
Given that Hallucination-Associated Neurons (H-Neurons) emerge during pre-training rather than instruction tuning or RLHF, what does this reveal about how hallucination-prone behaviour is encoded duri…
Self-improving Artificial Intelligence (AI) agent evaluation loop architecture
2026-03-05
What is the most principled architecture for a Self-Improving AI Agent Evaluation Loop — specifically, how should a nested inner/outer loop be designed so that a "Meta-Optimizer" rewrites system promp…
Evaluating and improving autonomous research loop quality
2026-03-03
How can the quality of research items produced by the `research-loop.yml` autonomous pipeline be systematically evaluated, and what changes to `research-prompt.md` and the loop's prompting strategy wo…
GitHub Specify, Ralph Loops, and Lisa Planning
2026-03-02
What is "Specify" in the context of GitHub-integrated AI development workflows, how does the Ralph loop implement proof-driven development in practice, and what role does Lisa planning play in the spe…