For this comparison, I kept Shampoo's exponent at its original value of 1/4. It is well-known that the modern "Shampoo^2" variant which uses 1/2 is more efficient, but this modification breaks the mathematical relationship to Muon and Spectral Descent, so I kept the original.
3/6
One motivation for me to add these optimizers was seeing @_arohan_, a senior researcher whose contributions I respect, repeatedly claim that Muon is Shampoo.
IIUC, his argument is that if we disable accumulation in both Muon and Shampoo, then they become...
4/6
...equivalent. This is correct. The problem is that Muon without accumulation is not Muon: It is Spectral Descent, which is >2x slower.
To go fast we need accumulation, and -- as shown in the figure -- the way it's added is what makes the difference between Muon and Shampoo.
5/6
Citations:
Shampoo is Gupta et al. (2018) and Anil et al. (2020)
Spectral descent is Carlson et al. (2015a) and Carlson et al. (2015b)
6/6
I've added two optimizers to the public benchmark:
(1) Shampoo (with its original 1/4 power).
(2) Spectral descent, which is equivalent to both Muon(mu=0) and Shampoo(b1=b2=0).
Result: Shampoo falls halfway between Muon & Adam; Spectral descent is ~2x slower.
Thread below
1/6Reproducible logs:
Shampoo:
Spectral descent:
As part of the public benchmark, further hyperparameter improvements are welcomed for any of these runs. All four use the same WSD lr schedule.
2/6For this comparison, I kept Shampoo's exponent at its original value of 1/4. It is well-known that the modern "Shampoo^2" variant which uses 1/2 is more efficient, but this modification breaks the mathematical relationship to Muon and Spectral Descent, so I kept the original.
3/6One motivation for me to add these optimizers was seeing @_arohan_, a senior researcher whose contributions I respect, repeatedly claim that Muon is Shampoo.
IIUC, his argument is that if we disable accumulation in both Muon and Shampoo, then they become...
4/6...equivalent. This is correct. The problem is that Muon without accumulation is not Muon: It is Spectral Descent, which is >2x slower.
To go fast we need accumulation, and -- as shown in the figure -- the way it's added is what makes the difference between Muon and Shampoo.
5/6Citations:
Shampoo is Gupta et al. (2018) and Anil et al. (2020)
Spectral descent is Carlson et al. (2015a) and Carlson et al. (2015b)
6/6
yes
I've added two optimizers to the public benchmark:
(1) Shampoo (with its original 1/4 power).
(2) Spectral descent, which is equivalent to both Muon(mu=0) and Shampoo(b1=b2=0).
Result: Shampoo falls halfway between Muon & Adam; Spectral descent is ~2x slower.
Thread below
1/6 ... Reproducible logs:
Shampoo:
Spectral descent:
As part of the public benchmark, further hyperparameter improvements are welcomed for any of these runs. All four use the same WSD lr schedule.
2/6 ... For this comparison, I kept Shampoo's exponent at its original value of 1/4. It is well-known that the modern "Shampoo^2" variant which uses 1/2 is more efficient, but this modification breaks the mathematical relationship to Muon and Spectral Descent, so I kept the original.
3/6 ... One motivation for me to add these optimizers was seeing @_arohan_, a senior researcher whose contributions I respect, repeatedly claim that Muon is Shampoo.
IIUC, his argument is that if we disable accumulation in both Muon and Shampoo, then they become...
4/6 ... ...equivalent. This is correct. The problem is that Muon without accumulation is not Muon: It is Spectral Descent, which is >2x slower.
To go fast we need accumulation, and -- as shown in the figure -- the way it's added is what makes the difference between Muon and Shampoo.
5/6 ... Citations:
Shampoo is Gupta et al. (2018) and Anil et al. (2020)
Spectral descent is Carlson et al. (2015a) and Carlson et al. (2015b)
6/6
Missing some Tweet in this thread? You can try to
Update