AI for Root Cause Analysis: A Lean Six Sigma Guide

AI root cause analysis

Artificial intelligence can examine large datasets, detect patterns and help improvement teams prioritise possible causes.  Those capabilities make AI potentially valuable for root cause analysis.

They also create a serious risk.

A model can identify a relationship without demonstrating why that relationship exists.

Effective root cause analysis therefore requires a combination of AI-supported analysis, process knowledge and disciplined validation.

For Lean Six Sigma practitioners, AI should help answer the question “Where should we investigate?” before the team answers the more demanding question “What evidence demonstrates that this is a cause?”

What is root cause analysis?

 

Root cause analysis is a structured attempt to identify the underlying factors contributing to a problem.

The distinction between symptom and cause matters.

Imagine a company experiencing a high number of late customer deliveries.

“Orders are waiting too long in the warehouse” describes part of the problem.

“Staff need to work faster” is an assumption.

A proper investigation asks what creates the delay.

Possible causes might include:

  • inaccurate inventory information
  • batch-processing rules
  • poor picking routes
  • equipment downtime
  • late order release
  • staffing patterns
  • priority changes
  • rework
  • transport cut-off times

The purpose of root cause analysis is to move from a broad list of possibilities towards causes supported by evidence.

Traditional root cause analysis tools

Lean Six Sigma practitioners can draw on several root cause analysis tools.

5 Whys

Repeatedly asking why can help a team move beyond an immediate symptom.

It is simple and useful, but it depends heavily on the team’s knowledge and should not be treated as proof by itself.

Fishbone diagram

A fishbone or cause-and-effect diagram helps organise possible causes into logical categories.

It is excellent for hypothesis generation.

Again, a cause appearing on the diagram does not make it true.

Pareto analysis

Pareto analysis helps identify categories contributing most strongly to a measured problem.

This can focus attention, but the largest category is not automatically the underlying root cause.

Process mapping

Mapping the process can reveal:

  • delays
  • handovers
  • loops
  • rework
  • unclear ownership
  • unnecessary activities

Statistical analysis

Suitable statistical methods can help practitioners evaluate relationships and differences using data.

AI adds another capability to this toolkit.

What AI changes

Traditional analysis can become difficult when a process produces hundreds of variables or millions of records.

AI and machine-learning methods can help identify patterns that would be difficult to discover manually.

Potential applications include:

  • defect prediction
  • equipment-failure analysis
  • complaint classification
  • anomaly detection
  • process-variable analysis
  • maintenance-log analysis
  • text analysis
  • multivariable pattern detection

Generative AI can also help teams organise information or propose hypotheses from supplied process descriptions.

But a plausible hypothesis is still a hypothesis.

Correlation is not causation

This is the central rule for AI-supported root cause analysis.

Suppose a model shows that products made during the night shift have a higher defect rate.

Possible conclusion:

The night shift causes defects.

Possible reality:

The night shift processes a higher proportion of difficult product variants.

Or perhaps preventive maintenance occurs immediately before the day shift.

Or experienced technicians are concentrated on the day shift.

Or environmental conditions differ overnight.

The relationship found by the model is useful because it tells the team where to investigate.

It does not automatically tell the team what to change.

A practical eight-step AI root cause analysis process

Step 1: Define the problem operationally

Avoid statements such as:

“Customers are unhappy.”

A better definition identifies what is happening and how it can be measured.

For example:

“The percentage of orders delivered more than 24 hours after the promised date increased from X to Y during the measured period.”

Use actual verified numbers when available.

Step 2: Establish process boundaries

Identify where the process starts and ends.

A SIPOC can help teams identify suppliers, inputs, major process stages, outputs and customers before detailed investigation.

Without boundaries, an AI dataset may combine information from several processes and produce misleading relationships.

Step 3: Map the actual process

Do not rely entirely on documented procedures.

Observe how work occurs.

Talk to people doing the work.

Identify:

  • queues
  • decisions
  • workarounds
  • exceptions
  • handovers
  • rework
  • informal rules

AI sees the data supplied to it. Process observation can reveal factors that were never recorded.

Step 4: Validate measurement

Before modelling, establish whether the data can answer the question.

Ask:

  • Are definitions consistent?
  • Are records complete?
  • Have collection methods changed?
  • Are timestamps reliable?
  • Are categories applied consistently?
  • Is the outcome variable measured correctly?

A model trained on unreliable operational data can produce a sophisticated explanation of a measurement problem.

Step 5: Generate hypotheses

Combine several sources:

  • process experts
  • frontline staff
  • historical information
  • fishbone analysis
  • 5 Whys
  • process maps
  • AI-supported exploration

The goal is broad thinking before narrowing.

Step 6: Use AI to prioritise patterns

Appropriate analytical systems may identify variables or combinations associated with the outcome.

For example, a model analysing service delays might identify:

  • case type
  • handover count
  • queue length
  • time of arrival
  • agent workload
  • approval route

as important predictors.

This helps prioritise investigation.

Step 7: Validate suspected causes

Now move beyond prediction.

Depending on the question, practitioners may use:

  • direct observation
  • stratification
  • statistical testing
  • regression
  • controlled experiments
  • designed experiments
  • pilot changes
  • additional data collection

A suspected cause becomes more credible when changing it produces the expected effect under suitable conditions.

Step 8: Confirm improvement and control

If removing or changing the suspected cause improves the process, establish controls to sustain the result.

Continue monitoring because the process can change.

Worked example: service request delays

Consider an organisation experiencing long resolution times for customer requests.

The team defines the problem and extracts operational data.

An AI model identifies the number of departmental handovers as one of the strongest predictors of long resolution time.

That does not prove handovers cause the delay.

The team maps high-handover cases.

It discovers that one request category lacks clear ownership. Requests repeatedly move between two departments because each interprets the responsibility differently.

The team tests a revised ownership rule.

Average handovers fall and resolution performance improves.

The sequence matters:

AI pattern → process investigation → mechanism identified → solution tested → result measured

That is stronger than:

AI pattern → immediate solution

Using generative AI during root cause analysis

Generative AI can help with several supporting activities.

A practitioner might ask a model to:

  • categorise incident descriptions
  • summarise maintenance notes
  • generate possible fishbone categories
  • identify questions for process interviews
  • organise workshop output
  • suggest additional hypotheses
  • explain statistical results in simpler language

Do not ask the model to “tell me the root cause” and treat the answer as evidence.

The model does not have access to unrecorded process reality unless the team provides it.

Generative systems can produce plausible but incorrect information.

That means Lean Six Sigma practitioners should apply quality thinking to AI outputs.

Possible defect categories might include:

  • invented facts
  • incorrect classifications
  • omitted information
  • inconsistent answers
  • unsupported recommendations
  • incorrect calculations

These outputs can be measured and analysed like other process outputs.

The organisation should decide what error level is acceptable for the intended application and when human review is required.

Explainability matters

For high-impact process decisions, knowing that a model predicts an outcome may not be enough.

The team may need to understand which factors influence the prediction and whether that relationship makes operational sense.

This is particularly important when:

  • employees are affected
  • customers are treated differently
  • safety is involved
  • regulated processes are involved
  • expensive changes are proposed

NIST’s AI risk-management resources emphasise managing AI risks across the lifecycle rather than assuming technical performance alone makes a system trustworthy.

Lean Six Sigma practitioners can contribute to that effort by insisting on measurable requirements, validation and control.

Common mistakes in AI root cause analysis

  • Treating feature importance as proof: A variable being important to a model does not prove it causes the outcome.
  • Ignoring process knowledge: Models can detect relationships that make no operational sense.
  • Analysing dirty data: Poor data quality can produce misleading patterns.
  • Missing variables: The true driver may never have been recorded.
  • Using historical bias as future policy: Past operational decisions can be embedded in historical datasets.
  • Implementing before testing: A recommendation should be evaluated before large-scale deployment where practical.