Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway

Published in ICML, 2026

Authors: Hee-Sung Kim, Sungyoon Lee

Venue: International Conference on Machine Learning (ICML), 2026

We show that the implicit bias of large-step Gradient Descent toward flat minima overrides the architectural bias of multi-pathway deep linear networks, reversing the “winner-takes-all” symmetry breaking predicted by Gradient Flow. Instead of collapsing onto a single dominant pathway, the network is driven into a re-balancing phase at the Edge of Stability that restores shared representations across parallel pathways.