skillrank_skill shipshitdev/evaluationMIT · open registry

skill detail

← registry

Evaluation

shipshitdev/evaluation

Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choices, or measuring improvements over time.

communityprovisionaltestingbuildevaluationframeworks

skillrank score

███░░░░░░░░░░░░░20

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

31

shipshitdev/skills

bundles/ai-agents/skills/evaluation

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install shipshitdev/evaluation
$ curl -fsSL skillrank.dev | sh