How does it do on CIFAR-10, or even better, ImageNet?
Lerc 1 hours ago [-]
It might be beneficial while not being optimal on its own.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
[1]: https://arxiv.org/abs/2605.31022
How does it do on CIFAR-10, or even better, ImageNet?
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.