Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway
Published in ICML, 2026
Authors: Hee-Sung Kim, Sungyoon Lee
Venue: International Conference on Machine Learning (ICML), 2026
We show that the implicit bias of large-step Gradient Descent toward flat minima overrides the architectural bias of multi-pathway deep linear networks, reversing the “winner-takes-all” symmetry breaking predicted by Gradient Flow. Instead of collapsing onto a single dominant pathway, the network is driven into a re-balancing phase at the Edge of Stability that restores shared representations across parallel pathways.
Download Paper | arXiv | Poster | Slides |
