A · Physics of AI 25–35 min

Broken scale symmetries in undercomplete linear autoencoders

Farhad Pashakhanloo, Jacob A. Zavatone-Veth · 2026-10-06
5 min
  1. 理解 scale symmetry
  2. 看 SGD drift 的方向
  3. 理解 edge of stability 为什么截断漂移
20 min
  1. solution manifold geometry
  2. fast/slow timescale separation
  3. effective scale dynamics
Deep dive
  1. Noether-like conserved quantity
  2. stochastic reduction
  3. 与 deep-network progressive sharpening 的联系

Big picture

这篇论文讨论一个非常 Physics-of-AI 的问题:

如果网络输出的函数已经完全不变,训练为什么还会继续改变参数?

linear autoencoder 的最优函数是 PCA projection,但参数化不是唯一的。

如果 encoder 与 decoder 写成 We,WdW_e, W_d,那么在合适条件下:

We→aWe,Wd→a−1WdW_e \to a W_e, \qquad W_d \to a^{-1} W_d

可以保持组合函数 WdWeW_dW_e 不变。

于是 global minima 不是单独一个点,而是一条带有 scale symmetry 的 solution manifold。

Continuous dynamics: symmetry and conservation

在连续 gradient flow 中,这种 scale symmetry 对应一个 Noether-like conserved quantity。直觉上:

如果只看连续优化方程,故事到这里就结束了。

What SGD changes

真实训练使用:

作者发现,在 PCA solution manifold 附近,encoder / decoder 两侧的 stochastic geometry 并不对称。

某些方向上的 minibatch noise 会被 loss curvature 和参数化几何“整流”,于是原本看似无偏的 stochastic fluctuations 变成有方向的慢漂移。

结果是:

这是一种 broken scale symmetry,不是“模型继续学到了新的 input-output mapping”。

Fast and slow variables

论文最漂亮的地方之一是 timescale separation:

因此可以把 fast relaxation 积掉,得到 scale variable 的 effective stochastic dynamics。

这种“fast mode → eliminate → slow effective theory”的思路与统计物理非常相似。

Edge of stability

漂移不会无限继续。

decoder scale 增大时,Hessian 的最大特征值也上升。finite-step gradient descent 的稳定性条件最终接近:

η μmax∼2\eta\,\mu_{\mathrm{max}} \sim 2

其中 η\eta 是 learning rate,μmax\mu_{\mathrm{max}} 是最大 Hessian eigenvalue。

于是 scale drift 被 edge of stability 截断。

这给 progressive sharpening 提供了一个很有意思的解释:

sharpness 可以继续增加,即使模型函数早已基本不变。

因此 sharpness 并不天然是 parameterization-independent 的“学习质量”指标。

Authors claim vs interpretation

Authors claim

在可解的 undercomplete linear autoencoder 中,有限步长 minibatch SGD 系统性破坏 scale symmetry,产生沿 functionally equivalent solution manifold 的定向 drift,并由 stability boundary 限制最终 scale。

Interpretation

这篇论文的价值主要不是 linear autoencoder 本身,而是提供一种研究深度学习动力学的范式:

symmetry → conserved quantity → perturbation breaks symmetry → timescale separation → effective slow dynamics → stability boundary

这套语言和凝聚态 / 统计物理训练非常兼容。

Caveats

Historical Reading Chain

1989 · Baldi & Hornik — Neural Networks and Principal Component Analysis
linear autoencoder 与 PCA solution geometry 的经典起点。
Read ↗
2021 · Kunin et al. — Neural Mechanics
最值得先读:用 symmetry 与 broken conservation laws 描述 deep-learning dynamics 的直接理论语言来源。
Read ↗
2021 · Cohen et al. — Gradient Descent at the Edge of Stability
理解为什么 finite-step stability boundary 会决定训练后期状态。
Read ↗
2023 · Pashakhanloo & Koulakov — SGD-Induced Drift of Representation
同一作者关于 minimum-loss manifold 上 stochastic drift 的直接前作。
Read ↗

Reading checklist

Sources and attribution