$mu$LO: Compute-Efficient Meta-Generalization of Learned Optimizers

2026-03-19 19:00 GMT · 4 months ago aimagpro.com

arXiv:2406.00153v5 Announce Type: replace
Abstract: Learned optimizers (LOs) have the potential to significantly reduce the wall-clock training time of neural networks. However, they can struggle to optimize unseen tasks (meta-generalize), especially when training networks wider than those seen during meta-training. To address this, we derive the Maximal Update Parametrization ($mu$P) for two state-of-the-art learned optimizer architectures and propose a simple meta-training recipe for $mu$-parameterized LOs ($mu$LOs). Our empirical evaluation demonstrates that LOs meta-trained with our recipe substantially improve meta-generalization to wider unseen tasks when compared to LOs trained under standard parametrization (SP) using the same compute budget. We also empirically observe that $mu$LOs exhibit unexpectedly improved meta-generalization to deeper networks ($5times$ meta-training) and surprising generalization to much longer training horizons ($25times$ meta-training) when compared to SP LOs.