What is FMEA? A Complete Guide to Measuring and Reducing Risk

Head scratch

Most organisations are good at reacting to things that have already gone wrong. Far fewer are good at working out what is likely to go wrong next, and doing something about it before it happens.

FMEA, or Failure Mode and Effects Analysis, is a structured method for doing exactly that. It gives you a repeatable way to list the ways something could fail, judge how bad each failure would be, and rank them so you can spend your limited time and money on the risks that actually matter.

This guide assumes you have never measured risk before. We will start with why risk measurement is worth doing at all, look at where FMEA came from, then walk through how to run one step by step with a worked example you can follow at your kitchen table.

Why bother measuring risk at all?

Before we get into method, it is worth being clear about the problem risk measurement solves.

Risk is probability multiplied by impact

The working definition is simple. The risk attached to an event is how likely it is to happen, multiplied by how much damage it does if it does happen.

That definition does useful work immediately. It tells you that a catastrophic event that almost never happens and a trivial event that happens constantly can carry similar amounts of risk. It also tells you that you cannot judge a risk sensibly by looking at likelihood or severity alone, which is exactly what most people do instinctively.

Your instincts about risk are unreliable

Human beings are famously bad at estimating risk. We over-weight events that are vivid, recent or heavily reported, and under-weight events that are slow, mundane or familiar. People worry more about dramatic and unlikely dangers than about the ordinary ones that are far more likely to hurt them.

This matters at work because the risks that get attention tend to be the ones that are easiest to picture, not the ones most likely to cause damage. Without a structured method, the loudest voice in the room sets the priorities.

Comparing is easier than estimating

Here is the insight that makes FMEA practical. Asking someone to state the absolute probability of a failure is asking for a number they do not have. Asking them whether this failure is more or less likely than that one is a question they can usually answer well.

FMEA is built on relative judgement rather than absolute measurement. You are not trying to calculate a true probability. You are trying to sort a list into a sensible order so that the things at the top get dealt with first. That is a much lower bar, and it is achievable with the knowledge your team already has.

Where FMEA came from

Head scratch

FMEA is not a modern management fad. It is one of the oldest systematic techniques for analysing failure, and it has roughly seventy-five years of use behind it.

The method was first formalised by the United States military in 1949, in a procedure called MIL-P-1629, Procedures for Performing a Failure Mode, Effects and Criticality Analysis. The problem it was built to solve was reliability in weapons systems: engineers needed a way to work out what would happen when a component failed, and to classify those failures by their effect on the mission and on personnel safety.

In the 1960s NASA adopted the approach for the Apollo programme. When there is no possibility of recall and no second attempt, predicting failure in advance stops being good practice and becomes the only option. FMEA went on to be used across later NASA missions and was picked up by civil aviation for the same reason.

The automotive industry followed in the mid-1970s. Ford introduced FMEA in response to the safety and reputational problems surrounding the Pinto, and developed its own handbook covering concept, design and process analyses. Other manufacturers in the United States and Europe followed.

In 1993 the Automotive Industry Action Group, formed by Ford, General Motors and Chrysler, published the first common FMEA reference manual, which is where the familiar severity, occurrence and detection scoring became standard. German manufacturers had developed a parallel approach under the VDA, and for years global suppliers had to satisfy two different methods. In 2019 the two bodies published a harmonised AIAG-VDA handbook, which reorganised the method into seven steps and introduced a different way of prioritising action.

Meanwhile the original military line continued. MIL-P-1629 became MIL-STD-1629 and then MIL-STD-1629A, which was formally cancelled in 1998 but is still widely referenced in defence and aerospace work. The international standard IEC 60812 covers the same territory, most recently updated in 2018.

The practical takeaway is that there is no single canonical FMEA. There is a family of closely related methods, and which one you use depends on your industry and who is asking. We cover the main variants in a separate guide to the types of FMEA.

The language of FMEA

Four terms do most of the work. Getting them straight early saves a lot of confused discussion later.

  • Failure mode is the way in which something fails. Not why it failed, and not what happened as a result, but the failure itself. "The cake is burnt" is a failure mode.
  • Effect is what the failure does to whoever or whatever is downstream. Burnt cake means a smoky kitchen and a disappointed family. Effects are usually described at more than one level: what it does locally, what it does to the wider system, and what the end user experiences.
  • Cause is the reason the failure mode occurred. The cake was left in the oven too long. Causes matter because you usually cannot act on a failure mode directly, only on its causes.
  • Control is whatever you already have in place that deals with the failure. Controls come in two flavours, and the difference is important: a prevention control stops the failure happening at all, while a detection control catches it after it has happened but before it reaches anyone who cares.

The distinction between prevention and detection is one of the most useful ideas in the whole method. Prevention reduces how often something goes wrong. Detection reduces how far the problem travels before someone notices. They are different jobs and they need different measures.

How to run an FMEA in eight steps

This is the working loop. It is deliberately simple, and it is the version worth learning first, whatever variant you eventually adopt.

1. Choose your subject

Decide precisely what you are analysing and where its boundaries lie. A product, a process, a service, a single piece of equipment. Being vague here is the most common way to waste a workshop, because the group ends up analysing different things at the same time. Write the scope down and agree it before anyone suggests a failure mode.

2. Identify the significant things that can fail

List the ways your subject could fail. Get the people who actually do the work into the room, because the value of the list depends entirely on real experience of things going wrong.

One rule matters more than any other at this stage: do not discard low-frequency events. The instinct to strike out anything that has only happened once is strong and it is wrong. Rare events with severe consequences are precisely the ones the method exists to surface.

3. Identify the effects of those failures

For each failure mode, describe what actually happens as a result. Be concrete. "Quality issue" tells you nothing. "Customer receives a unit that stops working within a week and requests a refund" gives you something you can score.

4. Score probability, severity and detectability

Score each failure mode on three scales, each running from 1 to 10. We cover exactly how these work in the next section.

5. Focus on the highest scores

Multiply the three scores to get a Risk Priority Number and sort your list. The top of that list is where your attention goes. The point of the exercise is to produce a ranking, not a set of precise measurements.

6. Take action to reduce the highest scores

Decide what you will change, who owns it and by when. Awareness of risk without action is worthless. An FMEA that produces a beautifully scored spreadsheet and no changes has failed.

7. Review

Re-score the failure modes you acted on and see what your changes actually bought you. This is also where you check whether the original scoring still looks sensible with hindsight.

8. Repeat until the residual risk is acceptable

Keep going round the loop until the organisation is comfortable with the level of risk that remains. Note the wording: the goal is not zero risk, which is unobtainable, but a residual level someone is prepared to sign their name against.

How the scoring works

Three scales, each from 1 to 10. The convention below is the one used in our training, and it is the most widely taught version.

Probability, where 10 means certain

How likely is this failure to occur? A score of 1 means it is close to unthinkable. A score of 10 means it is effectively certain to happen. Some variants call this occurrence.

Severity, where 10 means most severe

How bad is it when it happens? A score of 1 is an effect nobody would notice or care about. A score of 10 is catastrophic: serious injury, regulatory breach, loss of the business.

Detectability, the scale that runs backwards

This is the one that trips up almost every beginner, so it is worth stating plainly.

A failure that is easy to detect scores low. A failure that is hard to detect scores high.

The logic is consistent once you see it. All three scales are pointed so that a higher number means more risk, and an undetectable failure is more dangerous than an obvious one, because it travels further before anyone stops it. A defect that sets off an alarm the moment it occurs is far less threatening than one your customer discovers six months later.

Expect to explain this two or three times in a workshop. Expect at least one person to score it the wrong way round anyway. Building the direction into the wording of your scale, rather than relying on people to remember it, is the practical fix.

The Risk Priority Number

Multiply the three scores together and you get the Risk Priority Number:

RPN = Probability × Severity × Detectability

The result runs from 1 to 1000. Sort your failure modes by RPN, highest first, and you have a prioritised list of where to spend effort.

A worked example: baking a cake

Deliberately domestic, because the method is easier to learn when nobody is arguing about the subject matter. Two failure modes from a cake-baking FMEA:

Failure description How it fails Local effect Prob Sev Det RPN
In oven too long Cake burned because it is baked too long Burnt cake, smoke, unhappy family 6 7 7 294
Oven set to fan, not conventional Cake slightly overcooked, oven too hot Cake is a bit dry and not as nice to eat 6 4 7 168

Both failures are equally likely and equally hard to spot before the cake comes out. The difference in RPN comes entirely from severity: a burnt cake is a worse outcome than a slightly dry one. So the burnt cake gets dealt with first.

Notice how little arithmetic is involved, and how much of the value sits in the description columns. A well-written failure description is worth more than a precisely argued score.

Where RPN falls down

RPN is genuinely useful and genuinely flawed. You should know the flaws before you rely on it.

  • The scales are ordinal, but the maths treats them as if they were not. A severity of 8 is not twice as bad as a severity of 4 in any measurable sense, yet multiplying them implies that it is. The resulting number has no physical meaning.
  • Identical RPNs can hide very different risks. A score of 200 might be 10 × 5 × 4 or 5 × 5 × 8. The first involves a catastrophic outcome. The second does not. Treating them as equivalent is a serious mistake.
  • Thresholds get gamed. As soon as an organisation declares that anything above, say, 150 requires an action plan, scores start mysteriously landing at 144.
  • Low-probability, high-severity events get buried. A failure that would kill someone but has never yet happened scores low on probability, and the multiplication can push it below a dozen routine annoyances.

The last of these is serious enough that the 2019 AIAG-VDA handbook replaced RPN altogether with a system called Action Priority, which looks up each combination of scores in a table and returns High, Medium or Low, weighting severity first. Whether you need that depends on your industry, and we cover it in the variants guide.

For most organisations starting out, the sensible position is this: use RPN to sort your list, but never let the number make the decision on its own. Any failure mode with a severity of 9 or 10 gets reviewed regardless of where its RPN lands.

Taking action, which is the only part that matters

Head scratch

An FMEA that stops at scoring has achieved nothing. The scoring exists to direct action, and the analysis is only finished when something has actually changed.

For each failure mode you decide to tackle, record the mitigation you are putting in place, then re-score the same three dimensions assuming that mitigation is working. That gives you a revised RPN, and the difference between the two numbers, the RPN delta, tells you how much risk that change bought you.

This is more useful than it first appears. It lets you compare proposed actions against each other. An action costing very little that removes 200 points of RPN is a better use of the next hour than an expensive one that removes 40. It also gives you something concrete to show a sceptical manager, because "we reduced total RPN across the process by 38 per cent this quarter" is a defensible claim in a way that "we did a risk workshop" is not.

When you choose mitigations, remember the prevention and detection split. You have three levers, not one:

  • Make the failure less likely, by changing the design or the process. This is the strongest option and usually the hardest.
  • Make the consequences less severe, by containing the damage when it does occur. Often overlooked.
  • Make the failure easier to detect, by adding a check, an alarm or an inspection. Usually the cheapest and quickest, but it treats the symptom rather than the disease.

Teams gravitate towards detection because it is easy. A programme where every action is an extra inspection step is a programme that is accumulating cost without reducing the underlying failure rate. Watch for it.

Turning an FMEA into living risk KPIs

Head scratch

Here is where most FMEA programmes quietly die. The workshop is energising, the spreadsheet gets filled in, the actions get assigned, and eighteen months later nobody has looked at it. The analysis becomes a document rather than a tool.

The fix is to connect the FMEA to your ongoing measurement. Every failure mode you scored is a hypothesis about something that could go wrong, and most hypotheses can be monitored.

Two kinds of indicator are worth pulling out of an FMEA:

  • Lagging indicators count failures that have already happened. Incidents, escapes, complaints, cost of poor quality. They confirm whether your risk picture was right, but they tell you after the event.
  • Leading indicators track the conditions that make failure more likely, or the health of the controls that are supposed to catch it. Process drift, near misses, inspection pass rates, time taken to detect a problem, proportion of failure modes with no detection control at all.

Pair each significant risk with at least one of each. A lagging indicator on its own leaves you permanently reacting. A leading indicator on its own leaves you unable to tell whether your early warnings are actually predicting anything.

An indicator only earns its place if it has a defined formula, a threshold tied to how much risk you are willing to carry, a named owner, a measurement frequency and an agreed response for when it crosses the line. Without a predefined response it is not an early warning, it is just a number on a report.

It is also worth tracking the health of the FMEA programme itself. Useful measures include the number of high-scoring items still open, the age of overdue mitigation actions, the total RPN reduction achieved, the proportion of failure modes with no detection control, and how long it has been since each analysis was last reviewed. That last one is the single best predictor of whether your FMEA is still telling you the truth.

Keep the list short. A handful of well-chosen indicators per critical risk beats a dashboard nobody reads.

Common mistakes to avoid

  • Doing it alone. FMEA works best with a large amount of first-hand experience of failure. One person at a desk cannot supply that, however senior they are.
  • Striking out rare events. The failure that has only happened once is often the one that matters most.
  • Arguing about single points on the scale. If the room is split between 6 and 7, the answer almost never changes what you do next. Take the higher score and move on.
  • Vague failure descriptions. "Human error" is not a failure mode. It is a place where analysis stopped.
  • Treating the RPN as a measurement. It is a sorting device. Do not report it to two decimal places.
  • Scoring detectability backwards. Worth checking every single time.
  • Never reviewing it. A three-year-old FMEA describes a process that no longer exists.

When FMEA is the wrong tool

FMEA is one technique among many, and being honest about its limits will make you better at using it.

It works from the bottom up, one component or process step at a time, which makes it strong at finding single-point failures and weak at finding failures that emerge from several things going wrong together. If your concern is combinations, a fault tree analysis is a better fit. If you want to map the barriers on both sides of a single major hazard, a bowtie analysis will serve you better. For process plant hazards, HAZOP is the established choice. The international standard ISO 31010 sets out the wider family of risk assessment techniques if you want to see how they compare.

FMEA is also labour intensive and depends heavily on the knowledge in the room. Run it with people who have never seen the process fail and you will produce a confident-looking document describing risks that do not exist while missing the ones that do.

Use it when you have a definable subject, people with real experience of it, and a need to prioritise. That covers a great deal of ordinary business risk.

Where to go next

If you have followed this far, you know enough to run a basic FMEA on something small. The best way to learn it is to pick a process you understand well, get three or four colleagues in a room for ninety minutes, and work through the eight steps on a whiteboard. Your first attempt will be rough. That is fine, because you will repeat it.

Once you are comfortable with the basic method, the next question is usually which variant you should be using, particularly if a customer or regulator has opinions on the matter. That is covered in our companion guide to the types of FMEA, including DFMEA, PFMEA, FMECA and the AIAG-VDA method.

We work through FMEA in detail, alongside the wider KPI and measurement toolkit, on the KPI Black Belt programme.

Learn FMEA properly

FMEA is one of the techniques we cover in depth on the KPI Black Belt programme, alongside the rest of the measurement toolkit.

Find out about KPI Black Belt →

Useful Links