Endsley Model of Situational Awareness deep dive
- The Endsley model still anchors human factors work because it decomposes situational awareness into perception, comprehension, and projection, and later reviews treat that three-level structure as the dominant conceptual basis for measurement and interface evaluationSalmon et al. (2006)Endsley (2021)
- Human factors literature operationalises the model through several measurement families, especially SAGAT freeze probes, SPAM real-time probes, SART subjective ratings, and observer assessments, but review work recommends combining methods because no single measure captures complex team environments well enough on its ownSalmon et al. (2006)Endsley (2021)
- The strongest current direct-measure evidence appears to favour SAGAT, because Endsley's 2021 meta-analysis found it more sensitive than SPAM and free of several confounds that affect real-time probe methods, even though both predict performanceEndsley (2021)
- Recent AI-assisted decision-support evidence suggests that situational awareness improves oversight quality only when the interface and workflow sustain active verification, since error briefings and less aggregated evidence improve review quality more reliably than generic reminders of reviewer responsibilityKupfer et al. (2023)Conti (2025)
- Highly automated systems still require accurate situational awareness for safe decisions, and the technical challenge rises because context is assembled through sensor fusion, model inference, and rapidly changing machine behaviour rather than through a stable human-observable operating pictureIgnatious et al. (2023)Gao et al. (2023)
- The main theoretical limitation for oversight evaluation is that classical situational-awareness models are primarily individual-centred, while team and distributed-cognition critiques argue that effective oversight in complex socio-technical systems is distributed across people, artefacts, and coordination practicesGarbis (1998)Salmon et al. (2006)Gao et al. (2023)
- Official oversight guidance supports a limited-scope reading of the model, because regulators require understanding system limits and outputs while also requiring override rights, stop capability, competent staffing, monitoring, logs, independence, and manageable caseloadUnion (2024)Union (2024)Information (n.d.)National (2023)
- Compared with prior repository work, the evidence here suggests that queue pressure and weak governance often erode Level 2 comprehension and Level 3 projection before they eliminate nominal human sign-offGithub (n.d.)Github (n.d.)Kupfer et al. (2023)
Research Question
What is the Endsley Model of situational awareness, meaning the perception of relevant elements, comprehension of their meaning, and projection of their near-future status, how are its three levels defined and operationalised in human factors literature, and what current evidence exists on its usefulness and limitations for evaluating human oversight in highly automated and Artificial Intelligence (AI)-assisted systems?
Findings
Executive Summary
Endsley's three-level model remains a useful decomposition for evaluating human oversight because it defines situational awareness as perceiving relevant elements, comprehending their meaning, and projecting their near-future status before intervening in an automated system.
Current evidence does not support using the model as a complete oversight-evaluation framework in highly automated or AI-assisted settings, because modern failure modes also depend on workload, automation bias, reviewer authority, and coordination across people and artefacts.
Human factors measurement research supports using direct objective measures such as SAGAT, supplemented by workload, override-log, and review-quality metrics, rather than relying on a single situational-awareness score.
For modern AI oversight, the best-supported use of the Endsley model is as one component of a broader governance assessment that combines interface legibility, reviewer verification behaviour, caseload, and real stop-or-override power.
Key Findings
- The Endsley model still anchors human factors work because it decomposes situational awareness into perception, comprehension, and projection, and later reviews treat that three-level structure as the dominant conceptual basis for measurement and interface evaluation.
- Human factors literature operationalises the model through several measurement families, especially SAGAT freeze probes, SPAM real-time probes, SART subjective ratings, and observer assessments, but review work recommends combining methods because no single measure captures complex team environments well enough on its own.
- The strongest current direct-measure evidence appears to favour SAGAT, because Endsley's 2021 meta-analysis found it more sensitive than SPAM and free of several confounds that affect real-time probe methods, even though both predict performance.
- Recent AI-assisted decision-support evidence suggests that situational awareness improves oversight quality only when the interface and workflow sustain active verification, since error briefings and less aggregated evidence improve review quality more reliably than generic reminders of reviewer responsibility.
- Highly automated systems still require accurate situational awareness for safe decisions, and the technical challenge rises because context is assembled through sensor fusion, model inference, and rapidly changing machine behaviour rather than through a stable human-observable operating picture.
- The main theoretical limitation for oversight evaluation is that classical situational-awareness models are primarily individual-centred, while team and distributed-cognition critiques argue that effective oversight in complex socio-technical systems is distributed across people, artefacts, and coordination practices.
- Official oversight guidance supports a limited-scope reading of the model, because regulators require understanding system limits and outputs while also requiring override rights, stop capability, competent staffing, monitoring, logs, independence, and manageable caseload.
- Compared with prior repository work, the evidence here suggests that queue pressure and weak governance often erode Level 2 comprehension and Level 3 projection before they eliminate nominal human sign-off.
Assumptions
- Later accessible reviews quote the 1995 definition accurately enough that this item can analyse the model without a direct page-cited reading of the original article.
- Evidence from AI-assisted personnel selection, general human-AI decision support, and highly automated driving transfers to enterprise oversight because the common mechanism is human verification under automation, uncertainty, and workload.
- Recent human-AI teaming extensions are directionally useful for this item even though the accessible ATSA source is a preprint rather than a peer-reviewed journal article.
Analysis
The evidence clusters into model, measurement, behavioural, and governance strands that point in the same direction.
Within the model strand, the Endsley framework remains valuable because it makes oversight legibility testable, namely whether reviewers can see the right cues, understand them, and anticipate what will happen next if they do or do not intervene.
Measurement studies support direct probes such as SAGAT, yet the review literature also shows that team settings, dynamic queues, and real-time work require mixed measures because a single situational-awareness score can miss workload and coordination failures.
Recent automation-bias studies materially strengthen the case for broader oversight metrics, because better awareness cues only help when the workflow produces active checking instead of passive acceptance.
Official oversight sources complete the picture by adding authority, competence, staffing, logs, monitoring, and fallback, which makes the Endsley model most useful as a sub-framework inside a broader oversight assessment.
Risks, Gaps, and Uncertainties
- The foundational 1995 paper is represented here through later accessible quotations and reviews rather than direct page-cited use of the original article, so claims about subtle theoretical nuances should be treated cautiously.
- Recent human-AI situational-awareness extension work is thinner and less settled than the classical measurement literature, and one important accessible source in this item is a 2023 preprint rather than a peer-reviewed journal paper.
- Most direct objective measurement evidence comes from simulations or bounded experimental tasks rather than from live enterprise oversight queues, which limits external validity for large-scale production review environments.
Open Questions
- Which mixed measurement bundle best captures team-level situational awareness in production human-AI oversight queues without interrupting work?
- Can verification-intensity signals be instrumented cheaply enough to serve as a routine enterprise oversight metric rather than only an experimental construct?
- Which interface patterns most reliably improve Level 2 comprehension and Level 3 projection for reviewers supervising Large Language Model (LLM) systems?
sources
- [ ] Endsley (1995) Toward a Theory of Situation Awareness in Dynamic Systems - foundational paper identified; downstream definitional claims rely on accessible later quotations rather than direct page-cited use of the original article
- [x] Salmon et al. (2006) Situation awareness measurement: A review of applicability for command, control, communication, computers and intelligence (C4i) environments - accessible review quoting the canonical definition, summarising the three-level model, and comparing measurement approaches
- [x] Endsley (2021) A Systematic Review and Meta-Analysis of Direct Objective Measures of Situation Awareness: A Comparison of Situation Awareness Global Assessment Technique (SAGAT) and Situation Present Assessment Method (SPAM) - meta-analysis of 243 studies on direct objective situational-awareness measures
- [x] Artman and Garbis (1998) Team communication and coordination as distributed cognition - distributed-cognition critique of individual-centred situational-awareness models in team settings
- [x] Kupfer et al. (2023) Check the box! How to deal with automation bias in AI-based personnel selection - experiment linking verification intensity to better human review quality
- [x] Romeo and Conti (2025) Exploring automation bias in human-AI collaboration: a review and implications for explainable AI - review of workload, trust calibration, and blind reliance in human-AI decision-making
- [x] Gao et al. (2023) Agent Teaming Situation Awareness (ATSA): A Situation Awareness Framework for Human-AI Teaming - accessible framework paper extending situational-awareness concepts to human-AI teaming
- [x] Ignatious et al. (2023) Analyzing Factors Influencing Situation Awareness in Autonomous Vehicles, A Survey - recent survey showing why highly automated systems need accurate contextual awareness and why maintaining it is technically difficult
- [x] European Union (2024) AI Act Article 14 - official human-oversight requirements including automation-bias awareness, override, reverse, and stop capabilities
- [x] European Union (2024) AI Act Article 26 - official deployer duties covering competent human oversight, monitoring, suspension, and log retention
- [x] Information Commissioner's Office (ICO) (n.d.) Human review toolkit - official guidance on meaningful review, sampling, caseload, independence, and override logs
- [x] National Institute of Standards and Technology (NIST) (2023) Artificial Intelligence Risk Management Framework (AI RMF) Core - governance framework for ongoing monitoring, role clarity, and risk-proportionate review
| version | date | commit | summary |
|---|---|---|---|
| 1.0 | 2026-05-15 | 135d357 | Initial completion |