skillrank_skill hamelsmu/validate-evaluatorMIT · open registry

skill detail

← registry

Validate Evaluator

hamelsmu/validate-evaluator

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are deterministic;...

communityprovisionaldatacalibratellmjudge

skillrank score

███████░░░░░░░░░43

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

1.6k

hamelsmu/evals-skills

skills/validate-evaluator

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install hamelsmu/validate-evaluator
$ curl -fsSL skillrank.dev | sh