DSpark: Hardware-Aware Prefix Scheduling
Hard · Solved in PyTorch · Machine learning coding practice
Problem
Verifying a longer draft is not free. The target model scores every retained position in one causal pass, so each extra draft position adds a query row for every request in the batch, and a larger token batch runs fewer steps per second. Past some length the extra tokens a draft might win no longer pay for the slowdown they impose on the whole batch.
The scheduler in DSpark (DeepSeek-AI, 2026) resolves that trade-off per round by combining the draft head's confidence with speeds the serving engine measured at initialization.
Topics: inference, speculative-decoding, serving
The full statement, worked examples, hints and the test suite are available once you sign in. You can then solve DSpark: Hardware-Aware Prefix Scheduling in the browser and run it against the tests.