Why coordinate-level optimizers can behave differently on weights that represent the same model, and how quotient-aware updates remove the hidden gauge.
When Equivalent Weights Train Differently
jiha
Why coordinate-level optimizers can behave differently on weights that represent the same model, and how quotient-aware updates remove the hidden gauge.