skillrank_skill Prism-Shadow/agent-evaluationMIT · open registry

skill detail

← registry

Agent Evaluation

Prism-Shadow/agent-evaluation

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

communityprovisionaltestingrunonespecified

skillrank score

███████░░░░░░░░░41

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

1.1k

Prism-Shadow/penguin-harness

packages/skills/skills/agent-evaluation

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install Prism-Shadow/agent-evaluation
$ curl -fsSL skillrank.dev | sh