A large suite can run green and still miss the defect that matters. Test effectiveness asks what signal was produced, which risk it covers, what it cannot detect, and how to validate the gap.

Why this Skill is needed

The Test Effectiveness Analysis Skill connects test signals with defects and risks, missed-detection limits, false signals, and validation plans. It turns a metric or result into an evidence-bounded improvement conversation.

What the Skill does

Effectiveness is a relationship between a signal and a risk, not a count of tests.

Use this Skill to review a test plan, result, or metric and understand whether the supplied evidence supports the intended risk. It does not replace risk acceptance or turn a pass rate into a product-quality claim.

A good analysis names the test signal, risk coverage, missed-detection limit, false-signal concern, evidence state, validation method, owner, close condition, and residual risk.

Use it when a team questions whether more tests improve detection, a defect escaped despite green results, or a flaky signal is distorting a quality decision.

What it does not do

The Test Effectiveness Analysis Skill can organize material, evidence, and next actions for Signal judgment. It does not replace:

  • Confirmation of rules, scope, and risk by the accountable domain owner.
  • A real environment, account, dataset, log, or run record; static analysis does not become runtime evidence by itself.
  • Authorization for release, compliance, production actions, or residual-risk acceptance.
  • A Human decision when supplied sources conflict.

What it checks

The value of this Skill is not another keyword list. It connects each focus area to an observable input, a judgment, and a way to close the loop. Start with a small matrix based on the source Skill’s output contract:

FocusQuestion before analysisHandoff output
Signal judgmentWhat the test result actually indicatesSource and evidence state
Risk coverageWhich failure mode is protected or notDirect basis and limit
Missed-detection or false signalWhat green or red may fail to tell usUncertainty
Validation planSmallest check that can change the conclusionOwner and close condition

If a row has only a conventional expectation and no source or validation method, keep it open instead of turning it into a pass.

Audit inputs before you start

Before analyzing Signal judgment, classify the input into six evidence states. A gap is not automatically a failure, but it must not disappear inside the conclusion.

StateMeaningHow this Skill should handle it
knownDirectly supported by the supplied materialKeep the source, version, and time with the judgment
missingNeeded for this pass but not suppliedName the smallest evidence action and limit the conclusion
conflictingSources disagreeShow both sources and route the conflict to an owner
stalePresent but outside the relevant version or time windowMark freshness; old evidence is not current proof
out_of_scopeRelated but excluded from this passKeep the boundary explicit
assumptionsTemporarily adopted to continue analysisState how and when the assumption will be checked

Keep the input version, scope, environment, evidence locations, and accountable owner together. Without a run record, deliver analysis, design, or a validation plan—not an execution pass.

From problem to structured Finding

Connect the source, scope, evidence state, analysis, owner, action, close condition, and validation before writing the conclusion. The case below keeps this Skill’s identifier and domain context.

Keep the decision layers separate

For Signal judgment, do not compress four different kinds of language into “recommended to pass”:

LayerHow to write itApplication here
FactWhat the supplied material directly showsCite the source, version, input, or run record for the focus
Evidence-backed InferenceWhat several facts support togetherShow the inference chain and retain uncertainty
RecommendationThe smallest next actionName the evidence, review, execution, or regression path
Human DecisionWhat an accountable person must decideLeave scope, risk acceptance, resources, and release meaning to the owner

A complete case

This case follows Input, Analysis, Finding, Decision, and Validation. When material is incomplete, keep missing, conflicting, or assumptions visible instead of turning them into a pass.

Input

MaterialWhat to provideWhat to do when it is missing
Test signalSuite result, mutation result, defect detection, alert, or metric definitionKeep signal meaning explicit
Risk and defect evidenceFailure mode, escaped defect, affected journey, and expected behaviorDo not infer coverage from names
Window and environmentVersions, test data, execution time, environment, and comparison cohortMark missing context
Validation planSmall experiment, evidence source, owner, and stop conditionSeparate recommendation from execution

Use a request like this:

Use the test-effectiveness-analysis Skill.

Task: Analyze whether checkout contract tests detect the refund defects seen in the last two releases.
Inputs: [test results, defect records, risk list, changed scope, versions, environment]
Scope: [refund API and downstream ledger]
Constraints: [no unsupported pass or causal claim]

Cover test signal, risk coverage, missed-detection limit, false signal, validation plan, evidence state, owner, and residual risk.

Analysis

Use the input, matrix, and evidence state to form the judgment before writing the Finding; keep missing material as a gap.

Finding

Focused example: Explain a green suite after a refund escape

Contract tests verified the refund endpoint schema, but the escaped defect affected ledger posting after a successful response. The Skill should distinguish API contract signal from end-to-end risk coverage, identify the missed-detection limit, and propose one traceable ledger assertion or replay. It should not call the whole suite ineffective from one escape.

The analysis relates supplied signals to supplied risks. It cannot prove causal effectiveness, zero missed defects, or release readiness without an appropriate validation window and evidence.

Example finding: turn one problem into a handoff

The field example below shows the recording pattern; it is not an execution result.

If the supplied material cannot prove that Signal judgment meets its contract, write the finding like this. It does not invent the missing rule or turn missing evidence into a failure.

FieldExample wording
Source and scopeRecord the requirement, version, environment, and the concrete object for Signal judgment
FindingThe condition or result for Signal judgment is not yet traceable to evidence
Evidence statemissing / assumptions; use conflicting when sources disagree
Impact and priorityName the affected user, journey, or delivery decision without inflating severity
Owner and Human decisionAsk the product, engineering, security, or test owner to confirm the rule and trade-off
Action and close conditionAdd the smallest missing evidence; close only when source, judgment, and owner can be reviewed
ValidationName one repeatable check, query, or run and retain the raw artifact

The point is to let the next person walk from the finding back to the source and run an action that can change the decision.

Decision

The accountable owner confirms the decision question and risk trade-off; the Skill does not make that choice.

Validation

Before closing the finding, run the stated validation and retain the raw artifact. Without an execution record, the status remains unverified.

How a Finding enters the next stage

Handoff output

Output fieldWhy it existsExample status
Signal judgmentWhat the test result actually indicatesSource and evidence state
Risk coverageWhich failure mode is protected or notDirect basis and limit
Missed-detection or false signalWhat green or red may fail to tell usUncertainty
Validation planSmallest check that can change the conclusionOwner and close condition

Every conclusion should point to a source, evidence state, and next action. If evidence is missing, use pending, blocked, unassessed, or NOT_SCORED instead of filling the gap with confidence.

Next-stage route

At minimum, hand off the source, evidence state, owner, close condition, and validation action; the next-stage conclusion remains bounded by the evidence state.

How to prepare better input

Provide risk and defect records, test intent, results with versions and environment, changed surface, false-failure data, and the decision the analysis will inform. Without a comparable window, keep causal claims unassessed.

A useful handoff includes input versions, scope, time window, evidence index, assumptions, Human decision boundary, and the smallest validation action.

Working with other Skills

Common traps

  1. Using pass rate or test count as an effectiveness score.
  2. Treating one escaped defect as proof that every related test has no value.
  3. Ignoring false positives that consume attention and hide the signal.

Install and invoke

npx skills add https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/test-effectiveness-analysis -g
Use the test-effectiveness-analysis Skill.
Include objective, scope, versions, evidence paths, constraints, and decision boundary.
Audit inputs first, preserve evidence states, and finish with owner, close condition, and validation method.

FAQ

Can effectiveness be calculated from one release?

Usually not as a broad causal claim. Use the release as a bounded observation and state what comparison or validation is missing.

Does a mutation score prove real defect detection?

No. It is one signal whose mutant design, mapping, and limitations need review.

The Skill is most useful when attached to one real project artifact and kept with its source evidence. Start narrow, validate the uncertain part, and expand only when evidence supports it.

References

Source Skill and execution contract

The complete execution contract lives in Test Effectiveness Analysis prompt. The source directory may also contain evaluation cases and supporting material.

Static plans, file presence, and dry runs keep their evidence state. They do not become runtime proof, an all-passed claim, or release approval.

Share