top of page

Confirmation Bias in Troubleshooting: Preventing Diagnostic Fixation in High Risk Operations

Sep 27
14 min read
Wide-angle view of a refinery control room alarm panel during an abnormal operating condition

The first explanation in a troubleshooting event often feels like progress. A transmitter has failed before. A pump has been noisy all week. A compressor trip looks just like the last nuisance shutdown. In a control room, on a platform, beside a haul truck, or inside an emergency command post, that first explanation can bring order to uncertainty.


It can also trap the team.


High-risk operations depend on fast interpretation of weak, incomplete, and sometimes contradictory signals. Maintenance technicians, process operators, engineers, supervisors, emergency responders, pilots, clinicians, and control-room personnel all face the same cognitive problem: once a plausible diagnosis forms, people tend to search for confirmation, discount conflicting evidence, and stop exploring alternatives too early.


This is not a character flaw. It is a predictable feature of human cognition under pressure. The HSE and reliability question is whether organizations design troubleshooting systems that expose the bias before it becomes a pathway to loss.


Why Experienced Teams Lock Onto the First Plausible Answer


Confirmation bias is the tendency to notice, seek, and give weight to information that supports an existing belief while discounting information that does not. In troubleshooting, it often appears after a team forms an early explanation:


  • “That pressure spike is probably a bad transmitter.”

  • “This is the same vibration issue we had last month.”

  • “The gas detector is acting up again.”

  • “The patient is anxious, not deteriorating.”

  • “The warning is nuisance logic, not a real condition.”


Closely related concepts matter in operational settings.


Diagnostic fixation occurs when attention narrows around one diagnosis even as evidence changes. Aviation accident investigators, healthcare researchers, and human factors specialists use similar language to describe crews or clinicians who become anchored to an initial frame.


Premature closure is common in medical diagnosis literature, especially in work by researchers such as Pat Croskerry. It describes ending the diagnostic process before sufficient verification. The case feels solved, so active inquiry stops.


Selective attention means people focus on some cues while missing others. Under load, attention is not a neutral recording device. It filters. Salient, familiar, recent, or expected signals receive priority.


Cognitive psychology has studied these dynamics for decades. Peter Wason’s work on confirmation-seeking showed that people often test rules by looking for supporting examples rather than falsifying ones. Daniel Kahneman and Amos Tversky’s research on heuristics and biases helped explain why intuitive judgments can be efficient but error-prone. Gary Klein’s work on recognition-primed decision-making showed why experts often make fast, pattern-based decisions in real operations. That ability is valuable, but it carries a known vulnerability: if the pattern match is wrong, confidence can arrive before verification.


In high-risk work, expertise does not remove bias. It changes its shape. Experienced personnel may be more skilled at detecting patterns, but they may also recognize the wrong pattern faster.


How Diagnostic Fixation Shows Up in High-Risk Operations


Diagnostic fixation rarely announces itself. It usually looks like competent people working hard on the wrong problem.


The Instrument Failure Assumption


Instrument failure is a reasonable hypothesis in many operations. Sensors fail. Impulse lines plug. Calibration drifts. Alarms chatter. Experienced operators know this.


The problem starts when “instrument failure” becomes the preferred explanation before the team has ruled out the hazardous process condition the instrument is warning about.


Process safety investigations have repeatedly shown how abnormal readings can be normalized or treated as instrumentation problems. The Buncefield fuel storage incident in the United Kingdom is a widely cited example from the UK Health and Safety Executive and associated investigation reports. During filling operations, level indication and independent high-level protection failed to prevent overfill. While Buncefield involved technical, organizational, and control-system failures, it also illustrates a core operational issue: when level information is unreliable or not acted on as a credible warning, inventory can continue moving while risk escalates.


A similar pattern can arise during startup, tank transfer, pipeline packing, well control, boiler firing, confined space monitoring, or marine cargo operations. A reading that appears “wrong” may be wrong because the instrument failed, or because the process has moved outside the team’s expected mental model.


A disciplined troubleshooting response treats an abnormal reading as both possibilities until evidence separates them.


Practical field test:


If the instrument is assumed bad, what independent evidence would prove the process is safe while the instrument is investigated?

That evidence may include redundant indication, field verification, mass balance, temperature trend, valve lineup confirmation, independent sampling, portable detection, or stopping the energy or material transfer while uncertainty is resolved.


The Familiar Fault Assumption


Maintenance teams often carry a history of recurring defects. They know which motor starter is unreliable, which hydraulic valve sticks, which compressor seal has been troublesome, and which conveyor belt tends to mis-track after rain.


That history is useful. It can also become a cognitive shortcut.


Consider a mining maintenance crew responding to repeated trips on a conveyor drive. Previous events involved a faulty proximity switch. The team changes the switch, resets the system, and restarts. The trip returns. At that point, the hypothesis should weaken. If the group keeps treating each recurrence as another control fault, it may miss a developing mechanical jam, seized idler, belt fire risk, overload condition, or structural obstruction.


In manufacturing, a robot cell that repeatedly faults on a guard-door interlock may indeed have a failing switch. It may also have an operator workaround, misaligned guarding, vibration-induced wiring damage, or unexpected motion causing a real protective stop. In utilities, a breaker that trips “again” may reflect nuisance behavior, or it may be responding correctly to an intermittent fault.


The warning sign is not that the team uses equipment history. The warning sign is that equipment history becomes stronger than fresh evidence.


The Emergency Response Assumption


Emergency responders also face confirmation traps. Early information is often fragmented and sometimes wrong. Dispatch reports, gas readings, visual cues, caller descriptions, weather, and site knowledge compete for attention.


A hazmat team may initially frame symptoms as a small chemical release because that is what the first report suggests. A marine response crew may treat smoke as an electrical cabinet fire while cargo heating, battery failure, or fuel vapor involvement develops. A confined space rescue may begin with the assumption of trauma when atmospheric hazard is the primary threat.


Emergency command systems are designed to manage this uncertainty, but fixation can still occur when the incident action plan hardens too soon. Good commanders keep the working diagnosis visible, provisional, and open to challenge.


What Accident Investigations Teach About Fixation


Accident reports rarely use the phrase “confirmation bias” as the sole explanation. Good investigations avoid reducing complex events to one mental error. They look at equipment design, procedures, training, staffing, supervision, communications, maintenance, management systems, and regulatory controls.


Still, many official investigations describe patterns consistent with diagnostic fixation.


Aviation Shows the Cost of a Wrong Frame


Aviation is valuable to HSE because it has a mature investigation culture and strong human factors research. Accident reports from bodies such as the National Transportation Safety Board, Bureau d’Enquêtes et d’Analyses, and aviation regulators show how crews can misinterpret unreliable or unexpected indications.


Air France Flight 447, investigated by the BEA, is often discussed in human factors training because the crew faced unreliable airspeed indications at high altitude and did not maintain a shared, accurate understanding of the aircraft state. The event was not simply “confirmation bias,” and it involved automation mode awareness, training, startle, communication, and manual flying at altitude. Yet it shows how quickly crews can become absorbed in one understanding of the problem while the aircraft’s actual energy state deteriorates.


The aviation lesson for industrial operations is direct. During abnormal conditions, teams need methods that keep the actual system state visible. They need to ask what the plant, machine, vessel, or aircraft is doing physically, not only what the control system appears to be saying.


Healthcare Shows the Risk of Premature Closure


Healthcare diagnostic safety research has extensively examined premature closure. Clinicians operate under time pressure, incomplete information, interruptions, and high stakes. Studies and reviews by researchers such as Croskerry describe how anchoring, availability, confirmation bias, and search satisfaction can contribute to missed or delayed diagnoses.


The relevance to HSE is not that troubleshooting a pump is the same as diagnosing a patient. The relevance is the structure of the cognitive task. Both require professionals to interpret uncertain signals, revise hypotheses, and avoid declaring the problem solved too early.


Healthcare has responded with practices that translate well to high-risk operations:


  • Diagnostic checklists for high-risk presentations

  • Second opinions in uncertain cases

  • Cognitive forcing strategies that require alternative diagnoses

  • Time-outs when the clinical picture does not match the assumed condition

  • Escalation when response to treatment is not as expected


Industrial troubleshooting needs the same discipline when abnormal conditions carry major accident potential.


Process Safety Shows How Normalization Builds the Trap


Major accident investigations by organizations such as the U.S. Chemical Safety and Hazard Investigation Board and UK HSE often show that warning signs existed before loss events. The signs may have been ambiguous, intermittent, or normalized by prior experience.


The Deepwater Horizon investigations, including work by the U.S. government and the CSB, describe critical moments where well integrity information and test results were misinterpreted. The event involved a complex mix of technical, organizational, contractor interface, design, and decision factors. One relevant lesson is that ambiguous test data in a high-hazard system must be treated with structured skepticism. A desired interpretation can become very compelling when schedule pressure, prior experience, and authority gradients are present.


The BP Texas City refinery explosion investigation by the CSB also shows the danger of abnormal startup conditions, instrumentation and alarm issues, procedure weaknesses, and organizational factors. Again, it should not be reduced to a single bias. Yet the broader lesson holds: during startup and abnormal operations, teams are vulnerable to interpreting deviations through familiar or expected frames while energy and inventory accumulate.


Why Conventional HSE Systems Can Miss the Problem


Most HSE management systems require procedures, competency, risk assessments, permit controls, audits, incident reporting, and management review. These are necessary. They do not automatically create good diagnostic behavior in live troubleshooting.


Several gaps are common.


Procedures Describe Tasks, Not Reasoning


Maintenance and operating procedures often tell people what to do, not how to reason when the results do not make sense. A procedure may include steps for calibration, reset, isolation, or restart. It may not ask:


  • What else could explain this symptom?

  • What evidence would disprove the current diagnosis?

  • Which hazards increase if this diagnosis is wrong?

  • What condition requires stopping work or escalating?


When procedures omit diagnostic gates, experienced teams fill the gap using memory, informal norms, and local practice.


Investigations Focus on the Final Error


Incident reviews sometimes identify “failure to recognize” or “operator error” without examining why a particular interpretation made sense at the time. That misses the design of the trap.


A better investigation asks:


  • What cues were available?

  • Which cues were ambiguous or unreliable?

  • What histories or prior events shaped expectations?

  • What information was missing?

  • What pressures favored a quick diagnosis?

  • How easy was it to get independent verification?

  • Did systems invite challenge or punish delay?


This shifts attention from blame to learning. It also produces stronger controls.


Metrics Reward Restoration More Than Diagnosis


Many organizations track downtime, mean time to repair, maintenance backlog, schedule adherence, and production loss. Those metrics matter. They can also send a quiet signal that fast restoration is the mark of competence.


If teams receive praise for quick resets but little recognition for stopping to verify a weak diagnosis, diagnostic discipline erodes. The organization may not say “restart quickly,” but the measurement system may imply it.


Useful reliability cultures value speed after clarity. They do not treat uncertainty as a nuisance to be suppressed.


Control Rooms Can Overload Attention


Alarm floods, poor human-machine interface design, nuisance alarms, and weak alarm rationalization increase the chance of selective attention. Standards and guidance such as ISA 18.2 and the Engineering Equipment and Materials Users Association guidance on alarm systems reflect a long-standing recognition that alarm systems must support diagnosis, not simply produce alerts.


When an operator faces too many alarms, they will prioritize. That prioritization may be skilled, but it is still vulnerable to expectation and recent experience. If the interface makes the true abnormal condition hard to see, bias has room to grow.


Warning Signs That a Team Is Becoming Fixated


Diagnostic fixation is easier to interrupt early. Supervisors, control-room leads, field engineers, maintenance planners, and incident commanders should listen for language and behavior that shows the team has stopped testing the diagnosis.


Common warning signs include:


  • The same explanation is repeated after new conflicting evidence appears.

  • People say “it always does this” without checking whether the current context matches prior events.

  • Abnormal readings are labeled “bad instruments” before independent verification.

  • The team changes parts repeatedly without explaining why the part would cause all symptoms.

  • The same reset or bypass is attempted more than once without escalation.

  • Dissenting observations are treated as distractions.

  • Communication narrows to the people who already agree.

  • The group pays attention to one parameter while related parameters drift.

  • The response does not work, but confidence in the diagnosis remains high.

  • A supervisor asks for status and receives activity updates rather than evidence updates.


One practical phrase helps expose fixation:


“What would we expect to see if our diagnosis is wrong?”


That question changes the nature of the discussion. It moves the team from defending a conclusion to testing it.


Practical Countermeasures for HSE and Reliability Teams


Bias cannot be trained away by telling people to “be aware.” Awareness helps, but high-risk operations need designed countermeasures. The goal is not slower troubleshooting. The goal is more reliable troubleshooting under uncertainty.


Use Diagnostic Time-Outs at Defined Triggers


A diagnostic time-out is a short pause to reassess the working diagnosis, evidence, risk, and next action. It should be brief, operational, and triggered by pre-agreed conditions.


Good triggers include:


  • The first corrective action fails.

  • Two or more indicators conflict.

  • A safety-critical instrument is assumed faulty.

  • A trip, alarm, or shutdown is reset more than once.

  • The plant is in startup, shutdown, emergency operation, or degraded mode.

  • The consequence of being wrong includes loss of containment, uncontrolled energy, fire, explosion, collapse, radiation exposure, or life safety risk.

  • Field conditions do not match the control-room picture.


A useful time-out format can be completed in less than two minutes:


  1. What is our current diagnosis?

  2. What evidence supports it?

  3. What evidence conflicts with it?

  4. What are the top two alternate explanations?

  5. What is the worst credible consequence if we are wrong?

  6. What must we verify before restart, reset, re-entry, or continued operation?


This is not bureaucracy. It is a stop point for reasoning.


Ask Disconfirming Evidence Questions


Most troubleshooting naturally asks, “What confirms this?” Safer troubleshooting also asks, “What would prove this wrong?”


Examples:


  • If this pressure transmitter is faulty, what should the downstream pressure, flow, and tank level show?

  • If this vibration is only a sensor problem, why did bearing temperature change?

  • If this gas detector is false alarming, what does the portable meter show at the source and downwind?

  • If this is a nuisance trip, why did the motor current rise before the trip?

  • If this is a familiar hydraulic fault, why did cycle time change on a different actuator?


Disconfirming questions should be built into permits, shift handovers, control-room logs, troubleshooting work packs, and abnormal operating procedures where the hazard warrants it.


Require Alternate Hypotheses for High-Consequence Faults


For critical systems, one diagnosis is not enough. Teams should identify at least two plausible alternatives before taking action that could increase risk.


A pump low-flow alarm might be:


  • Instrument failure

  • Suction blockage

  • Closed or partially closed valve

  • Cavitation

  • Mechanical failure

  • Process condition outside design assumptions


A gas alarm might be:


  • Faulty detector

  • Real leak near the sensor

  • Vapor migration from another area

  • Calibration gas or maintenance activity

  • Cross-sensitivity, where applicable

  • Ventilation change concentrating vapors


The team does not need to spend equal time on every hypothesis. It needs to keep enough alternatives alive until evidence rules them out.


Use Independent Review Without Creating Delay Theater


Independent review works best when it is specific. “Get a second opinion” is vague. Better triggers include:


  • Safety-critical instrument declared failed

  • Critical alarm suppressed or inhibited

  • Protective trip bypassed for troubleshooting

  • Restart after unexplained trip

  • Conflicting control-room and field indications

  • Abnormal condition during simultaneous operations


The independent reviewer should not be told the preferred answer first. If possible, give them the raw symptoms and ask for their assessment. This reduces anchoring.


In practice, this may involve a shift supervisor, control-room lead, discipline engineer, senior technician, process safety engineer, marine superintendent, incident safety officer, or OEM specialist. The key is independence from the first diagnostic frame.


Structure Troubleshooting for Critical Equipment


Structured troubleshooting does not mean every fault needs a long checklist. It means critical faults receive proportional discipline.


For high-risk equipment and operations, structured methods may include:


  • Fault trees or cause maps for recurring critical failures

  • Decision tables for shutdown, continue, or escalate

  • Alarm response procedures linked to process consequences

  • Verification steps before reset or restart

  • Required cross-checks between field and control room

  • Defined limits for repeated resets

  • Temporary management of change for bypasses, inhibits, or degraded safeguards

  • Post-event review when the initial diagnosis was wrong or nearly wrong


Reliability engineering methods such as root cause analysis, failure modes and effects analysis, layers of protection analysis, and bow-tie analysis can support this work. The HSE value comes when those analyses influence live decisions, not when they sit in documents.


Train With Scenarios That Include Misleading Cues


Training often presents clean faults with clear symptoms. Real events are messier. Scenario design should include misleading but realistic cues:


  • A known unreliable sensor gives an alarm during a real release.

  • A familiar pump fault occurs at the same time as a suction restriction.

  • A nuisance trip happens during startup, but temperatures trend abnormally.

  • A gas reading appears near a detector with a maintenance history, but wind and ventilation explain the pattern.

  • A control valve problem masks an upstream blockage.


Aviation simulator training has long used surprise, ambiguity, and crew resource management concepts to improve response to abnormal events. Industrial control-room simulators, emergency exercises, tabletop drills, and field fault simulations can use the same principle. The point is not to trick people. It is to practice changing the diagnosis when the evidence changes.


Leadership Questions That Change the Quality of Troubleshooting


Senior leaders and operational managers do not need to become cognitive psychologists. They do need to shape the conditions in which diagnosis happens.


Useful leadership questions include:


  • Where do we allow repeat resets without formal escalation?

  • Which instruments do crews often mistrust, and what is being done about that mistrust?

  • During abnormal operations, how do teams prove the system is safe before continuing?

  • What are the top recurring faults that may be masking different failure modes?

  • Do our procedures require disconfirming evidence for high-consequence assumptions?

  • How often do post-incident reviews find that early signs were present but discounted?

  • Are alarm floods and nuisance alarms teaching operators to ignore the system?

  • Do supervisors reward careful verification, or only rapid restoration?

  • When field and control-room information conflict, who has authority to stop the operation?

  • Are engineers and maintainers trained to explain uncertainty clearly during live events?


Those questions are practical because confirmation bias in troubleshooting is partly a management system issue. People reason inside conditions created by design, staffing, training, interface quality, supervision, and production expectations.


What Good Looks Like in the Field


A mature troubleshooting culture does not treat every initial diagnosis as suspect. That would be inefficient and frustrating. Experienced judgment remains essential.


The difference is that good teams make the diagnosis testable.


A strong field response sounds like this:


  • “Our working diagnosis is a faulty level transmitter.”

  • “The evidence supporting that is the flatline trend and prior drift history.”

  • “The evidence against it is the rising inlet flow and no confirmed outlet flow.”

  • “The alternate hypotheses are actual level increase or outlet restriction.”

  • “We are stopping transfer until the field gauge and independent high-level check are verified.”

  • “If the independent reading confirms rising level, we shift to overfill prevention response.”


That exchange is short. It preserves speed. It also makes thinking visible.


In a maintenance context, the same discipline might sound like this:


  • “We replaced the proximity switch and the fault returned.”

  • “That weakens the original diagnosis.”

  • “Before another reset, we will inspect mechanical travel and motor current trend.”

  • “If we cannot explain all symptoms, we escalate to engineering before restart.”


For emergency response:


  • “The initial report suggests a small ammonia leak.”

  • “Readings downwind are higher than expected for that source.”

  • “We are expanding isolation and checking for a second release point.”

  • “Entry is paused until atmospheric data matches the plan.”


These examples show the operational value of humility. Not vague humility, but engineered humility expressed through verification, escalation, and disciplined doubt.


Professional References and Further Reading


  • U.S. Chemical Safety and Hazard Investigation Board Major accident investigation reports and safety videos, including refinery, chemical processing, and offshore events.


  • UK Health and Safety Executive Guidance and investigation material on major hazards, human factors, alarm management, and process safety.


  • National Transportation Safety Board and Bureau d’Enquêtes et d’Analyses Aviation accident investigation reports with extensive human factors analysis.


  • NASA Human Factors and Aviation Safety Research Research and operational lessons on decision-making, crew coordination, automation, and abnormal situations.


  • Pat Croskerry and Diagnostic Safety Literature Peer-reviewed work on cognitive forcing strategies, premature closure, anchoring, and diagnostic error in healthcare.


  • Daniel Kahneman, Amos Tversky, and Cognitive Bias Research Foundational research on heuristics, judgment under uncertainty, and decision-making.


  • Energy Institute Human Factors Guidance Practical resources for managing human and organizational factors in major hazard industries.


  • ISA 18.2 and EEMUA Alarm Management Guidance Widely used references for alarm system lifecycle management and control-room performance.


Professional Takeaway


Confirmation bias during troubleshooting is not a soft issue. It is a reliability and process safety risk that appears when capable people face uncertainty, time pressure, familiar faults, and imperfect information.


The control is not to tell teams to “try harder” or “pay attention.” The control is to design troubleshooting so early explanations stay provisional until tested. Use diagnostic time-outs. Ask for disconfirming evidence. Require alternate hypotheses for high-consequence faults. Bring in independent review before resets, bypasses, restarts, or continued operation under uncertainty.


In high-risk operations, the safest diagnosis is not the first plausible answer. It is the one that survives disciplined challenge.


bottom of page