skill detail
← registryComparative Evaluation
Owl-Listener/comparative-evaluation
A/B testing, side-by-side comparison, and preference ranking for AI outputs.
skillrank score
29
SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.
source
★ 157
Owl-Listener/ai-design-skills
claude-plugin/evaluation/skills/comparative-evaluation
open on GitHub ▸eval status
eval pending
Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.
install
$ skillrank install Owl-Listener/comparative-evaluation$ curl -fsSL skillrank.dev | sh