Skip to main content
Promotional banner for the pentest readiness checklist
Category: Risk Assessment & Analysis

Failure Mode and Effects Analysis

Also known as: FMEA, Failure Modes and Effects Analysis, Failure Mode & Effects Analysis
Simply put

Failure Mode and Effects Analysis (FMEA) is a structured, step-by-step method for identifying the ways a product, design, or process might fail and understanding the effects of those failures. It is used to prioritize which potential failures matter most so that teams can correct problems proactively rather than reacting to them after they occur. The goal is typically to improve the reliability and safety of complex systems by reducing the severity or likelihood of failures.

Formal definition

FMEA is a systematic technique for identifying potential failure modes within a design, product, or process, analyzing their causes and effects, and prioritizing them for mitigation. In classic FMEA practice, each identified failure mode is commonly evaluated against rating scales for Severity (the seriousness of the effect), Occurrence (the likelihood the cause or failure arises), and Detection (the likelihood the failure would be detected before reaching the customer or causing harm); these ratings are frequently combined into a Risk Priority Number (RPN = Severity × Occurrence × Detection) to help rank failures for attention. Mitigation is then typically structured around reducing the severity of a failure mode or effect, or lowering its probability of occurrence, so that processes can be corrected proactively. The specific rating scales, RPN methodology, and terminology vary across editions of standards and industry guidance, and some contemporary approaches supplement or replace RPN with alternative prioritization methods; practitioners should confirm the applicable framework and scoring conventions against the primary source relevant to their sector.

Why it matters

FMEA matters because it shifts risk management from a reactive posture to a proactive one. Rather than waiting for a product, design, or process to fail and then responding to the adverse event, teams use FMEA to identify potential failure modes in advance, understand their effects, and correct problems before they reach customers or cause harm. In complex systems, where the interaction of many components creates numerous ways for something to go wrong, this structured foresight can be the difference between a controlled design improvement and a costly or safety-critical failure.

The method's value lies in prioritization. Not every potential failure warrants equal attention, and resources for mitigation are finite. By evaluating failure modes against consistent criteria and ranking them, FMEA helps teams direct effort toward the failures that carry the greatest severity, are most likely to occur, or are hardest to detect before they cause damage. This makes it a widely used tool for improving the reliability and safety of complex systems across sectors such as manufacturing, healthcare, and engineering.

It is worth noting that FMEA is generally a leading-practice technique rather than a universal binding legal requirement, though it may be mandated or expected within particular sectors, standards, or contractual arrangements. Its effectiveness depends heavily on the quality of the input from those conducting the analysis, and rankings such as the Risk Priority Number are relative aids to judgment, not guarantees that a given failure has been eliminated. Practitioners should confirm the scoring conventions and framework applicable to their sector against the primary source.

Who it's relevant to

Risk Managers
FMEA gives risk managers a systematic way to identify potential failure modes and prioritize them for treatment, supporting proactive risk reduction rather than reaction to events after they occur. It complements broader risk assessment activities by providing a structured, repeatable ranking method for design and process risks.
Quality and Reliability Engineers
In manufacturing and engineering contexts, FMEA is a core tool for improving the reliability and safety of complex systems. Engineers use it to surface potential failures within a product design or process, analyze causes and effects, and direct corrective effort toward the failures that matter most.
Healthcare and Patient Safety Teams
Healthcare organizations use FMEA to evaluate clinical and operational processes for possible failures and to correct them proactively rather than reacting to adverse events. This makes it relevant to teams focused on patient safety and process improvement in complex care environments.
Internal Auditors and Compliance Professionals
While FMEA is primarily a risk management technique, auditors and compliance professionals may encounter it when assessing whether an organization has adequately identified and mitigated process and design risks. Understanding its scoring conventions and limitations helps in evaluating the rigor of an organization's proactive risk controls, though applicability varies by sector and framework.

Inside FMEA

Failure Mode
A specific manner in which a process, product, system, or component could fail to perform its intended function. FMEA analyzes each identified failure mode individually rather than in aggregate.
Effects of Failure
The consequences that a given failure mode may have on the process, downstream operations, the end product, or the customer. A single failure mode may carry multiple effects of differing significance.
Causes of Failure
The underlying conditions or mechanisms that could trigger a failure mode. Identifying causes supports the design of targeted preventive measures rather than symptom-level responses.
Severity (S) Rating
A numeric rating, typically on a defined scale, estimating the seriousness of the effect of a failure should it occur. Higher ratings indicate more serious consequences.
Occurrence (O) Rating
A numeric rating estimating the likelihood or frequency with which a given cause and its associated failure mode are expected to arise, typically on a defined scale.
Detection (D) Rating
A numeric rating estimating the ability of existing controls to detect a failure mode or its cause before impact is realized. In common conventions, a higher rating reflects lower detectability (i.e., a greater chance the failure escapes detection); practitioners should confirm scale direction against the specific standard or template in use.
Risk Priority Number (RPN)
In classic FMEA, a composite score commonly calculated as RPN = Severity × Occurrence × Detection, used to help prioritize which failure modes warrant attention. Some newer approaches supplement or replace RPN with action-priority methods, so treatment of RPN varies across editions and guides.
Current Controls
Existing prevention or detection measures already in place for a failure mode or cause. These inform the Occurrence and Detection ratings and provide the baseline against which improvements are assessed.
Recommended Actions and Follow-up
Proposed measures to reduce severity, occurrence, or improve detection, along with assigned ownership, target dates, and often a reassessment of ratings after actions are implemented.

Common questions

Answers to the questions practitioners most commonly ask about FMEA.

Is FMEA a compliance control that eliminates the risk of failures?
No. FMEA is an analytical and risk assessment technique, not a control that eliminates risk. It is used to systematically identify potential failure modes, their causes, and their effects so that they can be prioritized and treated. In risk terms, FMEA supports the identification and assessment of risk; it does not by itself modify or eliminate risk. Any reduction in risk comes from the corrective actions or controls implemented as a result of the analysis, and even then risk is typically reduced rather than eliminated. FMEA also does not, on its own, establish compliance with any particular law or standard, though it may be used to support a broader compliance or quality management program.
Does a low Risk Priority Number (RPN) mean a failure mode can be safely ignored?
Not necessarily. In classic FMEA, the RPN is typically calculated as Severity × Occurrence × Detection, where each factor is rated on a defined scale. A low composite RPN can be misleading because it may combine a very high Severity with low Occurrence and Detection ratings, masking a failure mode with serious consequences. For this reason, many practitioners and guides recommend treating high Severity ratings as a trigger for action regardless of the overall RPN, and some approaches supplement or replace RPN with action-priority methods. RPN is best understood as one prioritization aid rather than a definitive threshold for whether a failure mode warrants attention.
What are the core elements typically scored in a classic FMEA?
Classic FMEA typically scores each identified failure mode against three factors: Severity (the seriousness of the effect of the failure), Occurrence (the likelihood that the cause of the failure will arise), and Detection (the likelihood that existing controls will detect the failure or its cause before it reaches the customer or causes harm). Each is rated on a defined scale, and in many standards and industry guides these are multiplied to produce the Risk Priority Number (RPN = Severity × Occurrence × Detection). The specific rating scales, definitions, and prioritization conventions vary across editions, sectors, and organizations, so the applicable reference should be confirmed against the primary source in use.
How should an organization decide which failure modes to act on first?
Prioritization commonly draws on the Severity, Occurrence, and Detection ratings, whether combined into an RPN or evaluated through an action-priority approach. Many practitioners treat high Severity as warranting attention even where the composite score is low, and consider the combination of factors rather than a single number. Organizations often set their own thresholds or decision rules aligned to their risk appetite and tolerance. Because conventions differ across frameworks and editions, the prioritization method and any thresholds should be documented and applied consistently, and treated as a matter of informed judgment rather than a mechanical cutoff.
Who should typically be involved in conducting an FMEA?
FMEA is generally most effective as a cross-functional exercise. Depending on scope, participants often include those with knowledge of the design, process, or system under review, along with quality, engineering, operations, and, where relevant, risk or compliance functions. The intent is to bring together the range of expertise needed to identify credible failure modes, causes, effects, and existing controls, and to rate them consistently. The appropriate composition varies by the type of FMEA (for example, design versus process), the sector, and the organization's structure.
When should an FMEA be updated after it is first completed?
FMEA is often described as a living document rather than a one-time exercise. It is commonly revisited when there are relevant changes, such as to the design, process, operating environment, applicable requirements, or when new failures or near-misses come to light. Reviewing the analysis after corrective actions are implemented can help confirm whether the ratings still hold and whether residual risk is acceptable. The frequency and triggers for updates vary by organization and sector, and should be defined in the relevant procedure and verified against any applicable standard.

Common misconceptions

The Risk Priority Number provides an objective, absolute measure of risk that can be compared across different analyses.
RPN values depend on the rating scales, definitions, and judgment applied in a given FMEA, so they are typically used to prioritize within a single analysis rather than as portable or absolute risk measures. Because the three factors are multiplied, very different risk profiles can yield the same RPN, which is one reason some approaches supplement it with severity thresholds or action-priority logic.
FMEA is a compliance deliverable completed once and then filed away.
FMEA is generally treated as a living document intended to be revisited as designs, processes, controls, or conditions change. Its value depends on periodic review and on following through on recommended actions rather than on the initial documentation alone.
A low Occurrence rating means a failure mode can be disregarded.
A failure mode with low likelihood may still carry high severity, and prioritization typically weighs severity, occurrence, and detection together. Many practitioners give special attention to high-severity effects regardless of occurrence, since a low probability does not eliminate the potential consequence.

Best practices

Define and document the Severity, Occurrence, and Detection rating scales explicitly before scoring, including the direction of the Detection scale, so ratings are applied consistently and remain interpretable by later reviewers.
Conduct FMEA with a cross-functional team so that failure modes, effects, causes, and existing controls are captured from multiple operational and technical perspectives rather than a single viewpoint.
Avoid relying on the Risk Priority Number alone; consider severity thresholds or action-priority approaches so that high-consequence failure modes are not overlooked when their multiplied score appears moderate.
Assign clear ownership, target dates, and follow-up verification for each recommended action, and re-score the relevant ratings after actions are implemented to confirm the intended risk reduction.
Treat the FMEA as a living document and schedule reviews when designs, processes, controls, or operating conditions change, rather than completing it as a one-time exercise.
Confirm the specific FMEA methodology, template, and scale conventions against the primary standard or industry guide being followed, since terminology and the treatment of RPN evolve across editions.
a promotional banner asking how ready are you for PCI DSS 4.0? With a call-to-action to get the checklist now.