skillrank_skill bdiasti/eval-testingMIT · open registry

skill detail

← registry

Eval Testing

bdiasti/eval-testing

Build evaluation frameworks for AI agents with LLM-as-judge, rule-based evals, and golden datasets. Use when testing agents, evaluating RAG quality, or creating compliance benchmarks.

communityprovisionaltestingbuildevaluationframeworks

skillrank score

███░░░░░░░░░░░░░18

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

21

bdiasti/maestro-bundle-cli

templates/bundle-ai-agents/skills/eval-testing

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install bdiasti/eval-testing
$ curl -fsSL skillrank.dev | sh