Speculative decoding orchestrator for LLM serving: draft-model ensemble with rejection sampling, KV-cache-aware admission control and self-tuning draft lengths
golang performance orchestration developer-tools rejection-sampling serving kv-cache llm-inference speculative-decoding draft-models
-
Updated
Sep 12, 2026 - Go