Google publishes Switch Transformer, scaling a sparse model to 1.6T parameters
Switch Transformer routed each token to a single expert, scaling total sparse parameters to about 1.6 trillion without activating the full model for every token.
MODEL
Official links | Chat | Code Plan | Agent tools | API |
|---|---|---|---|---|
| Global |
English reports linked to this entity will appear here as their translations are published.