autorenew

SKILL DETAIL

LLM Hallucination Testing

Use this skill when you need evidence-bounded claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review; triggers include LLM 幻觉 and LLM hallucination.

StatusStable
TypeAtomic Skill
DomainSoftware testing
SDLCTest design
Good forQA / DEV
LanguageChinese / English
EvalsEvals ✓
Synced2026-09-15

Why this Skill

It turns this Skill's method into a quality input that can be executed, reviewed, and reused.

  • Use this skill when you need evidence-bounded analysis, design, or validation preparation for claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review.
  • Keep the analysis focused on claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review; do not replace business owners or Human risk acceptance, exception approval, or safety decisions.
  • Never invent system behavior, fields, model outputs, data, thresholds, root causes, execution records, or pass claims.
  • Static design, plans, file presence, or a dry run retain their evidence state and cannot become proof of real execution.
  • Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.

When to use

Use this Skill
  • Use this skill when you need evidence-bounded analysis, design, or validation preparation for claim-to-source relations, unsupported assertions, abstention, uncertainty, and evidence review.
  • Use it to review an Agent, RAG, or LLM plan, result, or evidence set and produce actionable improvements.
  • Use it when context is incomplete but a bounded first pass with assumptions, gaps, and human-decision boundaries is still useful.
Common pitfalls
  • Listing checks without triggers, expected concerns, owner roles, close conditions, and evidence.
  • Treating adjacent tests or model tools as a complete LLM hallucination judgment.
  • Using unexplained numbers for false precision or writing correlation as causation.
  • Refusing incomplete input, or pretending that incomplete evidence is conclusive.

Input

Minimum Input
  • The current task scope, objective, and subject under review.
Recommended Input
  • Project goal
  • Test scope
  • Constraints
Optional Context
  • Relevant code or configuration
  • Historical results
  • Logs and metrics

Output

The output follows this Skill's method and makes facts, assumptions, risks, and next steps explicit.

Judge the output value before installing

  1. 01Read and follow prompts/llm-hallucination-testing.md, including its input audit, domain coverage, and output order.
  2. 02Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to claim-to-source relation, unsupported assertion, abstention, uncertainty, evidence review.
  3. 03Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.
  4. 04Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.
View full output structure
  1. 05When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.

How It Works

  1. 01Read and follow prompts/llm-hallucination-testing.md, including its input audit, domain coverage, and output order.
  2. 02Extract scope, environment, version, time window, constraints, success criteria, and available evidence, with attention to claim-to-source relation, unsupported assertion, abstention, uncertainty, evidence review.
  3. 03Separate confirmed facts, evidence-backed inferences, candidate recommendations, and Human decisions before ranking by risk and evidence strength.
  4. 04Turn high-risk items into preconditions, steps, expected behavior or decision criteria, required evidence, and a validation method.
  5. 05When input is incomplete, deliver a bounded first pass, state unsupported conclusions, and never present static material as execution evidence.

Install & Quick Start

Install command / SHELL
npx skills add \
  https://github.com/naodeng/awesome-qa-skills/tree/main/skills/en/testing-types/llm-hallucination-testing
  -g
llm-hallucination-testing.prompt
@skill llm-hallucination-testing

Using the current project context, produce an actionable result following this Skill.

Additional context:
[Paste project context or requirement]