autorenew

SKILL DETAIL

LLM Consistency Testing

Use this skill when you need evidence-bounded repeat inputs, version/model/prompt factors, invariants, variance evidence, and comparison boundaries; triggers include LLM 一致性 and LLM consistency.

StatusStable
TypeAtomic Skill
DomainSoftware testing
SDLCTest design
Good forQA / DEV
LanguageChinese / English
EvalsEvals ✓
Synced2026-09-15

Why this Skill

It turns this Skill's method into a quality input that can be executed, reviewed, and reused.

  • Use this skill when you need evidence-bounded analysis, design, or validation preparation for repeat inputs, version/model/prompt factors, invariants, variance evidence, and comparison boundaries.
  • Keep the analysis focused on repeat inputs, version/model/prompt factors, invariants, variance evidence, and comparison boundaries; do not replace business owners or Human risk acceptance, exception approval, or safety decisions.
  • Never invent system behavior, fields, model outputs, data, thresholds, root causes, execution records, or pass claims.
  • Static design, plans, file presence, or a dry run retain their evidence state and cannot become proof of real execution.
  • Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.

When to use

Use this Skill
  • Use this skill when you need evidence-bounded analysis, design, or validation preparation for repeat inputs, version/model/prompt factors, invariants, variance evidence, and comparison boundaries.
  • Use it to review an Agent, RAG, or LLM plan, result, or evidence set and produce actionable improvements.
  • Use it when context is incomplete but a bounded first pass with assumptions, gaps, and human-decision boundaries is still useful.
Common pitfalls
  • Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.
  • Treating adjacent tests or model tools as a complete LLM consistency judgment.
  • Using unexplained numbers for false precision or writing correlation as causation.
  • Refusing incomplete input, or pretending that incomplete evidence is conclusive.

Input

Minimum Input
  • The current task scope, objective, and subject under review.
Recommended Input
  • Project goal
  • Test scope
  • Constraints
Optional Context
  • Relevant code or configuration
  • Historical results
  • Logs and metrics

Output

The output follows this Skill's method and makes facts, assumptions, risks, and next steps explicit.

Judge the output value before installing

  1. 01Read and follow prompts/llm-consistency-testing.md, including its input audit, domain coverage, and output order.
  2. 02Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to repeat input, version factor, model factor, invariant, variance evidence.
  3. 03Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.
  4. 04Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.
View full output structure
  1. 05When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.

How It Works

  1. 01Read and follow prompts/llm-consistency-testing.md, including its input audit, domain coverage, and output order.
  2. 02Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to repeat input, version factor, model factor, invariant, variance evidence.
  3. 03Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.
  4. 04Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.
  5. 05When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.

Install & Quick Start

Install command / SHELL
npx skills add \
  https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/llm-consistency-testing
  -g
llm-consistency-testing.prompt
@skill llm-consistency-testing

Using the current project context, produce an actionable result following this Skill.

Additional context:
[Paste project context or requirement]