A large suite can run green and still miss the defect that matters. Test effectiveness asks what signal was produced, which risk it covers, what it cannot detect, and how to validate the gap.
Why this Skill is needed
The Test Effectiveness Analysis Skill connects test signals with defects and risks, missed-detection limits, false signals, and validation plans. It turns a metric or result into an evidence-bounded improvement conversation.
What the Skill does
Effectiveness is a relationship between a signal and a risk, not a count of tests.
Use this Skill to review a test plan, result, or metric and understand whether the supplied evidence supports the intended risk. It does not replace risk acceptance or turn a pass rate into a product-quality claim.
A good analysis names the test signal, risk coverage, missed-detection limit, false-signal concern, evidence state, validation method, owner, close condition, and residual risk.
Use it when a team questions whether more tests improve detection, a defect escaped despite green results, or a flaky signal is distorting a quality decision.
What it does not do
The Test Effectiveness Analysis Skill can organize material, evidence, and next actions for Signal judgment. It does not replace:
- Confirmation of rules, scope, and risk by the accountable domain owner.
- A real environment, account, dataset, log, or run record; static analysis does not become runtime evidence by itself.
- Authorization for release, compliance, production actions, or residual-risk acceptance.
- A Human decision when supplied sources conflict.
What it checks
The value of this Skill is not another keyword list. It connects each focus area to an observable input, a judgment, and a way to close the loop. Start with a small matrix based on the source Skill’s output contract:
| Focus | Question before analysis | Handoff output |
|---|---|---|
| Signal judgment | What the test result actually indicates | Source and evidence state |
| Risk coverage | Which failure mode is protected or not | Direct basis and limit |
| Missed-detection or false signal | What green or red may fail to tell us | Uncertainty |
| Validation plan | Smallest check that can change the conclusion | Owner and close condition |
If a row has only a conventional expectation and no source or validation method, keep it open instead of turning it into a pass.
Audit inputs before you start
Before analyzing Signal judgment, classify the input into six evidence states. A gap is not automatically a failure, but it must not disappear inside the conclusion.
| State | Meaning | How this Skill should handle it |
|---|---|---|
| known | Directly supported by the supplied material | Keep the source, version, and time with the judgment |
| missing | Needed for this pass but not supplied | Name the smallest evidence action and limit the conclusion |
| conflicting | Sources disagree | Show both sources and route the conflict to an owner |
| stale | Present but outside the relevant version or time window | Mark freshness; old evidence is not current proof |
| out_of_scope | Related but excluded from this pass | Keep the boundary explicit |
| assumptions | Temporarily adopted to continue analysis | State how and when the assumption will be checked |
Keep the input version, scope, environment, evidence locations, and accountable owner together. Without a run record, deliver analysis, design, or a validation plan—not an execution pass.
From problem to structured Finding
Connect the source, scope, evidence state, analysis, owner, action, close condition, and validation before writing the conclusion. The case below keeps this Skill’s identifier and domain context.
Keep the decision layers separate
For Signal judgment, do not compress four different kinds of language into “recommended to pass”:
| Layer | How to write it | Application here |
|---|---|---|
| Fact | What the supplied material directly shows | Cite the source, version, input, or run record for the focus |
| Evidence-backed Inference | What several facts support together | Show the inference chain and retain uncertainty |
| Recommendation | The smallest next action | Name the evidence, review, execution, or regression path |
| Human Decision | What an accountable person must decide | Leave scope, risk acceptance, resources, and release meaning to the owner |
A complete case
This case follows Input, Analysis, Finding, Decision, and Validation. When material is incomplete, keep missing, conflicting, or assumptions visible instead of turning them into a pass.
Input
| Material | What to provide | What to do when it is missing |
|---|---|---|
| Test signal | Suite result, mutation result, defect detection, alert, or metric definition | Keep signal meaning explicit |
| Risk and defect evidence | Failure mode, escaped defect, affected journey, and expected behavior | Do not infer coverage from names |
| Window and environment | Versions, test data, execution time, environment, and comparison cohort | Mark missing context |
| Validation plan | Small experiment, evidence source, owner, and stop condition | Separate recommendation from execution |
Use a request like this:
Use the test-effectiveness-analysis Skill.
Task: Analyze whether checkout contract tests detect the refund defects seen in the last two releases.
Inputs: [test results, defect records, risk list, changed scope, versions, environment]
Scope: [refund API and downstream ledger]
Constraints: [no unsupported pass or causal claim]
Cover test signal, risk coverage, missed-detection limit, false signal, validation plan, evidence state, owner, and residual risk.
Analysis
Use the input, matrix, and evidence state to form the judgment before writing the Finding; keep missing material as a gap.
Finding
Focused example: Explain a green suite after a refund escape
Contract tests verified the refund endpoint schema, but the escaped defect affected ledger posting after a successful response. The Skill should distinguish API contract signal from end-to-end risk coverage, identify the missed-detection limit, and propose one traceable ledger assertion or replay. It should not call the whole suite ineffective from one escape.
The analysis relates supplied signals to supplied risks. It cannot prove causal effectiveness, zero missed defects, or release readiness without an appropriate validation window and evidence.
Example finding: turn one problem into a handoff
The field example below shows the recording pattern; it is not an execution result.
If the supplied material cannot prove that Signal judgment meets its contract, write the finding like this. It does not invent the missing rule or turn missing evidence into a failure.
| Field | Example wording |
|---|---|
| Source and scope | Record the requirement, version, environment, and the concrete object for Signal judgment |
| Finding | The condition or result for Signal judgment is not yet traceable to evidence |
| Evidence state | missing / assumptions; use conflicting when sources disagree |
| Impact and priority | Name the affected user, journey, or delivery decision without inflating severity |
| Owner and Human decision | Ask the product, engineering, security, or test owner to confirm the rule and trade-off |
| Action and close condition | Add the smallest missing evidence; close only when source, judgment, and owner can be reviewed |
| Validation | Name one repeatable check, query, or run and retain the raw artifact |
The point is to let the next person walk from the finding back to the source and run an action that can change the decision.
Decision
The accountable owner confirms the decision question and risk trade-off; the Skill does not make that choice.
Validation
Before closing the finding, run the stated validation and retain the raw artifact. Without an execution record, the status remains unverified.
How a Finding enters the next stage
Handoff output
| Output field | Why it exists | Example status |
|---|---|---|
| Signal judgment | What the test result actually indicates | Source and evidence state |
| Risk coverage | Which failure mode is protected or not | Direct basis and limit |
| Missed-detection or false signal | What green or red may fail to tell us | Uncertainty |
| Validation plan | Smallest check that can change the conclusion | Owner and close condition |
Every conclusion should point to a source, evidence state, and next action. If evidence is missing, use pending, blocked, unassessed, or NOT_SCORED instead of filling the gap with confidence.
Next-stage route
At minimum, hand off the source, evidence state, owner, close condition, and validation action; the next-stage conclusion remains bounded by the evidence state.
How to prepare better input
Provide risk and defect records, test intent, results with versions and environment, changed surface, false-failure data, and the decision the analysis will inform. Without a comparable window, keep causal claims unassessed.
A useful handoff includes input versions, scope, time window, evidence index, assumptions, Human decision boundary, and the smallest validation action.
Working with other Skills
- Test Gap Analysis:turns missed obligations into explicit gaps.
- Risk-Based Testing:connects risk evidence to test depth.
- Test Suite Health Analysis:adds suite-level reliability and maintenance evidence.
Common traps
- Using pass rate or test count as an effectiveness score.
- Treating one escaped defect as proof that every related test has no value.
- Ignoring false positives that consume attention and hide the signal.
Install and invoke
npx skills add https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/test-effectiveness-analysis -g
Use the test-effectiveness-analysis Skill.
Include objective, scope, versions, evidence paths, constraints, and decision boundary.
Audit inputs first, preserve evidence states, and finish with owner, close condition, and validation method.
FAQ
Can effectiveness be calculated from one release?
Usually not as a broad causal claim. Use the release as a bounded observation and state what comparison or validation is missing.
Does a mutation score prove real defect detection?
No. It is one signal whose mutant design, mapping, and limitations need review.
The Skill is most useful when attached to one real project artifact and kept with its source evidence. Start narrow, validate the uncertain part, and expand only when evidence supports it.
References
Source Skill and execution contract
The complete execution contract lives in Test Effectiveness Analysis prompt. The source directory may also contain evaluation cases and supporting material.
Static plans, file presence, and dry runs keep their evidence state. They do not become runtime proof, an all-passed claim, or release approval.
Reference links
- Test Effectiveness Analysis prompt:https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/test-effectiveness-analysis/prompts/test-effectiveness-analysis.md
- Test Effectiveness Analysis Skill source:https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/test-effectiveness-analysis
- Test Effectiveness Analysis details:https://inaodeng.com/en/qaskills/test-effectiveness-analysis/
- Awesome QA Skills on GitHub:https://github.com/naodeng/awesome-qa-skills