Gated DeltaNet: Projections and Head Shapes

Easy · Solved in PyTorch · Machine learning coding practice

Problem

Implement the independent input branches of Gated Delta Networks: Improving Mamba2 with Delta Rule (Yang, Kautz and Hatamizadeh, 2024). This exercise uses the same number of query, key, and value heads; convolution, normalization, and gate activations come later.

x has shape (B,T,D). weights has exactly six matrices in output-by-input layout: q, k have shape (H d_k,D); v, z have shape (H d_v,D); a, b have shape (H,D). num_heads is H and beta_bias has shape (H,). Compute xW^\mathsf{T} independently for every branch, adding beta_bias only to b.

Topics: gated-deltanet, linear-attention, llm-internals

The full statement, worked examples, hints and the test suite are available once you sign in. You can then solve Gated DeltaNet: Projections and Head Shapes in the browser and run it against the tests.

Related problems