skillrank_skill modular/eval-modelMIT · open registry

skill detail

← registry

Eval Model

modular/eval-model

Measures the task accuracy of text models served by MAX using standard benchmarks such as GSM8K, MMLU, HellaSwag, ARC, AIME, GPQA, TruthfulQA, WinoGrande, and BABILong. Use when benchmarking a served model, comparing it with model-card or reference scores, verifying that a...

communityprovisionalaimeasurestaskaccuracy

skillrank score

████░░░░░░░░░░░░28

SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.

source

125

modular/skills

eval-model

open on GitHub ▸

eval status

eval pending

Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.

install

$ skillrank install modular/eval-model
$ curl -fsSL skillrank.dev | sh