Muon^p: Muon with Fractional Spectral Powers
Yihe Dong ⋅ Will Sawin
Abstract
Muon replaces a gradient $G=USV^\top$ with its polar factor $UV^\top$, thereby flattening the singular spectrum, but full flattening discards singular-value information that may matter for adaptation. We introduce PowerMuon (Muon$^p$), a Muon-style optimizer that instead uses fractional spectral-power updates $US^pV^\top$ for rational $p\in(0,1)$, interpolating between Muon and gradient descent. To make it practical, we prove that fractional spectral powers cannot be computed by any fixed univariate polynomial iteration and derive low-degree odd bivariate recurrences that approximate $US^pV^\top$ using only matrix multiplications, preserving Muon’s matrix-multiplication-only structure and compute complexity. Empirically, Muon$^p$ is especially effective for finetuning: on billion-scale models, Muon$^p$ significantly improves validation perplexity and downstream performance. We further analyze when Muon$^p$ is less suitable, through the lens of spectral geometry.
Successful Page Load