Local SGD cadence as a Master-Stability-Function problem: call for a collaborator with synchronization-theory depth
📰 Reddit r/deeplearning
I've been working on a heuristic for when to AllReduce in heterogeneous Local SGD, one that's empirically battle-tested across six architecture families (MLP, LeNet, ResNet-20, char-RNN, GPT-nano, conv AE). On the He et al. 2015 ResNet-20 CIFAR-10 setup (published paper 91.25%, 200 epochs), an RTX 5060 Ti + GTX 1060 mix reaches 92.42%, above the published number, in less wall time than the 5060 Ti alone (91.66%). The heuristic watches ||pre-AllReduc
DeepCamp AI