Small-vocabulary MTP heads for memory-bound speculative decoding
Can a compact, separately trained MTP head replace the inherited full-width block and 248k-row output projection without surrendering the acceptance that makes speculation useful? The audit answered: no—function cuts pay an off-distribution tax that fidelity cuts do not, and the zero-training trimmed-vocabulary recipe is the one that survives.