skillrank_skill wshobson/llm-evaluationMIT · open registry

skill detail

← registry

llm-evaluation

wshobson/llm-evaluation

Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking.

communityprovisionalaiagent-skillsagentic-ai

skillrank score

██████████░░░░░░63

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

37.7k

wshobson/agents

plugins/llm-application-dev/skills/llm-evaluation

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install wshobson/llm-evaluation
$ curl -fsSL skillrank.dev | sh