skillrank_skill patricio0312rev/evaluation-harnessMIT · open registry

skill detail

← registry

Evaluation Harness

patricio0312rev/evaluation-harness

Builds repeatable evaluation systems with golden datasets, scoring rubrics, pass/fail thresholds, and regression reports. Use for "LLM evaluation", "testing AI systems", "quality assurance", or "model benchmarking".

communityprovisionaltestingbuildsrepeatableevaluation

skillrank score

████░░░░░░░░░░░░23

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

56

patricio0312rev/skills

ai-engineering/evaluation-harness

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install patricio0312rev/evaluation-harness
$ curl -fsSL skillrank.dev | sh