Code Quality · Version 1.0.0 · Reviewed 2026-08-02
Reviewer Quality Evaluator
Make a defensible decision about reviewer finding evaluation and severity calibration audit with evidence, explicit trade-offs, and a verification plan.
4 method steps
5 documented failure modes
4 diagnostic checks
7 quality gates
Scores reviewer output for detection accuracy, severity calibration, actionability, scope discipline, and signal-to-noise. It grounds the decision in labeled changes, reviewer findings, accepted dispositions, severity definitions, and false-positive evidence and explicitly prevents rewarding finding volume or eloquence while incorrect and non-actionable comments burden maintainers.
₹299 one-time
Get this skill archive
What it checks first
Reviewer Quality Evaluator scores reviewer output for detection accuracy, severity calibration, actionability, scope discipline, and signal-to-noise. It grounds the decision in labeled changes, reviewer findings, accepted dispositions, severity definitions, and false-positive evidence and explicitly prevents rewarding finding volume or eloquence while incorrect and non-actionable comments burden maintainers. Use it when the work involves Reviewer finding evaluation, Severity calibration audit, Review signal measurement.
- Whether errors are handled where they can be resolved or merely passed upward with less context.
- Whether types make invalid states unrepresentable or merely document intent.
- Ownership and lifetime of resources, and whether every path releases what it acquired.
- Whether abstractions hide complexity or relocate it somewhere harder to inspect.
Example task
Input
Apply the reviewer quality evaluator to our current reviewer finding evaluation work. We need a concrete decision, bounded changes, and evidence that the result is correct.
Expected output
Start with labeled changes, reviewer findings, accepted dispositions, severity definitions, and false-positive evidence. The highest-risk failure is rewarding finding volume or eloquence while incorrect and non-actionable comments burden maintainers. Weight correctness and reachability first, then actionability, calibration, and concise evidence. Verify the result by double-labeling a representative sample and reporting agreement, precision, recall, and severity error.