autorenew

SKILL DETAIL

LLM Testing

Use this skill when you need to test LLM behavior, failure modes, and evidence-based quality boundaries; triggers include llm testing.

StatusStable
TypeAtomic Skill
DomainSoftware testing
SDLCTest design
Good forQA / DEV
LanguageChinese / English
EvalsEvals ✓
Synced2026-09-15

Why this Skill

It turns this Skill's method into a quality input that can be executed, reviewed, and reused.

  • Use this skill when you need to test LLM applications for task quality, robustness, safety, factuality, and version regressions.
  • avoid exact-string assertions alone
  • pin and record model settings
  • evaluate stochastic outputs with repetitions and distributions
  • Listing checks without preconditions, expected outcomes, or evidence.

When to use

Use this Skill
  • Use this skill when you need to test LLM applications for task quality, robustness, safety, factuality, and version regressions.
  • Use it to review an existing plan, result, or evidence set and produce actionable improvements.
  • Use it when context is incomplete but a bounded first pass is still valuable.
Common pitfalls
  • Listing checks without preconditions, expected outcomes, or evidence.
  • Marking everything high priority and avoiding tradeoffs.
  • Substituting tool names or generic theory for domain reasoning.
  • Refusing incomplete input, or pretending incomplete evidence supports certainty.

Input

Minimum Input
  • The current task scope, objective, and subject under review.
Recommended Input
  • Project goal
  • Test scope
  • Constraints
Optional Context
  • Relevant code or configuration
  • Historical results
  • Logs and metrics

Output

The output follows this Skill's method and makes facts, assumptions, risks, and next steps explicit.

Judge the output value before installing

  1. 01Read and follow prompts/llm-testing.md, including its input contract, execution rules, minimum coverage, and output order.
  2. 02Add only context that changes the decision: scope, environment, version, constraints, evidence, and success criteria.
  3. 03Audit the input, then separate confirmed facts, working assumptions, and open questions.
  4. 04Rank by risk and evidence strength, and produce an artifact that can be executed or reviewed directly.
View full output structure
  1. 05If information is missing, deliver a bounded first pass and state which conclusions remain unsupported.

How It Works

  1. 01Read and follow prompts/llm-testing.md, including its input contract, execution rules, minimum coverage, and output order.
  2. 02Add only context that changes the decision: scope, environment, version, constraints, evidence, and success criteria.
  3. 03Audit the input, then separate confirmed facts, working assumptions, and open questions.
  4. 04Rank by risk and evidence strength, and produce an artifact that can be executed or reviewed directly.
  5. 05If information is missing, deliver a bounded first pass and state which conclusions remain unsupported.

Install & Quick Start

Install command / SHELL
npx skills add \
  https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/llm-testing
  -g
llm-testing.prompt
@skill llm-testing

Using the current project context, produce an actionable result following this Skill.

Additional context:
[Paste project context or requirement]