GShard paper demonstrates a sparse-expert translation model with more than 600B parameters
Google researchers introduce GShard automatic sharding and use it to train a multilingual translation MoE with more than 600B total parameters on 2,048 TPU v3 accelerators.