skillrank_skill davila7/agent-evaluationMIT · open registry

skill detail

← registry

Agent Evaluation

davila7/agent-evaluation

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent...

communityprovisionaltestingtestingbenchmarkingllm

skillrank score

█████████░░░░░░░59

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

30.2k

davila7/claude-code-templates

cli-tool/components/skills/ai-research/agent-evaluation

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install davila7/agent-evaluation
$ curl -fsSL skillrank.dev | sh