skill detail
← registryInference Serving Topology
claude-dev-suite/inference-serving-topology
LLM/model inference serving architecture: the engine → serving → orchestration layering (vLLM/SGLang/TensorRT-LLM, Triton, KServe/Ray Serve), KV-cache & continuous batching, prefill-decode disaggregation, and scaling. Architect-level topology, not model training. USE WHEN:...
skillrank score
SkillRank score blends community stars, real usage, and our eval lift; provisional until a skill is evaluated -- so popularity alone can't reach the top tier.
source
★ 27
claude-dev-suite/claude-dev-suite
skills/ai-systems/inference-serving-topology
open on GitHub ▸eval status
eval pending
Success delta, token delta, and trial count are not available yet. No eval number is shown until this skill has measured results.
install
$ skillrank install claude-dev-suite/inference-serving-topology$ curl -fsSL skillrank.dev | sh