Flaky Test Analysis: Why do flaky tests fail intermittently?

An intermittent failure is more than an annoying red build. It erodes trust in CI and makes real regressions easier to ignore. The first job is turning randomness into classifiable evidence.

The Flaky Test Analysis Skill separates product instability from timing, data leakage, and environment noise, then directs the fix to the responsible boundary.

This guide organizes the practice around clear inputs, boundaries, and outputs, so the next person can act on the conclusion with confidence.

Flaky Test Analysis Skill: what it is for

Flaky Test Analysis is for work that needs a clear, handoff-ready testing judgment. It keeps project material, the basis for each decision, and the next action on the same trail—so a reader can see what to inspect before choosing how to execute and review it. This guide works through one concrete scenario and keeps human decision boundaries visible.

Start with the source Skill

The complete execution contract lives in Flaky Test Analysis prompt. The source directory also contains 3 evaluation cases for checking whether an output follows the contract.

The entry point calls out these constraints:

  • do not hide failures with blind retries
  • quarantine needs an owner and exit criteria
  • root-cause claims require reproducible evidence

Begin with project facts

Put the material you have on the table. Gaps may remain; their status needs to stay explicit.

MaterialWhat to provideWhat to do when it is missing
Goal and scopeAnalyze intermittent UI regression failures and separate product defects, synchronization faults, unstable environments, and dirty dataName journeys outside this pass
Version and environmentRequirement version, build, environment, time windowStay in design or analysis mode
EvidenceRequirements, interfaces, logs, metrics, traces, or defectsSeparate facts, assumptions, and open questions
Decision boundaryRisk approver and actions that are not authorizedName the owner and next step

Use a request like this:

Use the flaky-test-analysis Skill.

Task: Analyze intermittent UI regression failures and separate product defects, synchronization faults, unstable environments, and dirty data
Inputs: [versions, links, log paths, or reports]
Scope: [included and excluded objects]
Constraints: [time, data, permissions, compliance]

Audit the inputs first. Order results by risk and evidence strength. Label unsupported claims as assumptions and give a validation method.

Make the result usable by the next person

Output fieldWhy it existsExample status
Finding or judgmentDescribes observed behavior, difference, or riskConfirmed / Assumption / Open
BasisPoints to a version, log, trace, test, or requirementsource_id or link
ImpactExplains affected users, journeys, or release decisionP0, P1, or accepted residual risk
Next actionNames verification work and an ownerOwner, date, expected evidence

Do not write “passed” without a run record, query result, or source artifact. Static analysis and runtime proof are different things.

Run one focused pass

Start with a bounded pass—Analyze intermittent UI regression failures and separate product defects, synchronization faults, unstable environments, and dirty data. Put the input version, time window, and accountable owner in one place. Then link each judgment to an artifact. Finish with one validation action that can change the decision.

Preserve first-failure evidence on each rerun; retry rate does not replace root-cause classification. The handoff should include an evidence index, assumptions that still need checking, and an action the next person can run without reconstructing the conversation. Plain work. It holds up.

Run one focused pass

Start with a bounded pass—Analyze intermittent UI regression failures and separate product defects, synchronization faults, unstable environments, and dirty data. Put the input version, time window, and accountable owner in one place. Then link each judgment to an artifact. Finish with one validation action that can change the decision.

Preserve first-failure evidence on each rerun; retry rate does not replace root-cause classification. The handoff should include an evidence index, assumptions that still need checking, and an action the next person can run without reconstructing the conversation. Plain work. It holds up.

Advanced use: turn one analysis into a maintained mechanism

Preserve first-failure evidence on each rerun; retry rate does not replace root-cause classification.

Keep input versions and source IDs with every result. When requirements, code, environment, or data change, recompute only affected judgments and mark them changed, unchanged, or needs-review. Old conclusions are not new evidence.

A three-Skill chain

flaky-test-analysischange-impact-analysisregression-test-selection

HandoffPayloadReceiver check
Upstream to flaky-test-analysisSource versions, scope, risk, open itemsStaleness and conflicts
flaky-test-analysis to downstreamJudgments, evidence index, residual risk, tasksExecutability and ownership
Feedback to flaky-test-analysisRuns, defects, changed factsBaseline and regression scope

Hand over a summary, an evidence index, and locations for the source artifacts. That gives the receiver enough context and keeps the trail recoverable.

Team gates

GateCheckFailure action
flaky-test-analysis inputVersion, environment, sources, and ownerStop and list gaps
flaky-test-analysis artifactMaterial claims have basis, status, and impactReturn for evidence
flaky-test-analysis executionCommand, query, or verification path is repeatableClassify infrastructure or test issue
flaky-test-analysis decisionResidual risk has an accepter and dateDo not enter the next stage

Common traps

  1. Listing checks without input conditions, expected results, or evidence.
  2. Marking every finding high priority and removing the team’s ability to choose.
  3. Refusing to produce a bounded first pass, or presenting guesses as facts.
  4. Treating one success or one anomaly as long-term behavior while ignoring repeated trials and version changes.

Two practical questions

Can I start with incomplete input?

Yes. Produce a constrained first pass with known facts, assumptions, gaps, and the smallest validation action. Missing environment, data, or permission cannot support an execution claim.

When is human confirmation required?

The accountable owner must confirm scope trade-offs, risk acceptance, production actions, data permission, and release decisions. The Skill organizes evidence and options; it does not grant authority.

Run Flaky Test Analysis with one real artifact and keep the input, output, human edits, and verification evidence in the same work chain. That is what makes the next change cheaper to assess.

References

Share