skill detail
← registryevaluating-llms-harness
Orchestra-Research/evaluating-llms-harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag).
skillrank score
55
SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.
source
★ 10.5k
Orchestra-Research/AI-Research-SKILLs
11-evaluation/lm-evaluation-harness
open on GitHub ▸eval status
eval pending
Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.
install
$ skillrank install Orchestra-Research/evaluating-llms-harness$ curl -fsSL skillrank.dev | sh