DSpark: Markov Draft Head

Medium · Solved in PyTorch · Machine learning coding practice

Problem

A speculative decoder is only as good as its drafter. Drafting K tokens with K sequential passes is slow, while K parallel heads are fast but blind to one another: each position guesses without seeing what the position before it chose, so the drafts are mutually incoherent and acceptance collapses after the first one.

Topics: inference, speculative-decoding, autoregressive

The full statement, worked examples, hints and the test suite are available once you sign in. You can then solve DSpark: Markov Draft Head in the browser and run it against the tests.

Related problems