🔬Ai2's Olmo-core 3 Trains MoEs 2.7x Faster Per GPU
TL;DR
Ai2 released Olmo-core 3, open training infrastructure for large mixture-of-experts models. A 47B MoE hit 52,000 tokens per second per GPU, versus 19,400 on the earlier stack.
Ai2 released Olmo-core 3, open training infrastructure for large mixture-of-experts models. A 47B MoE hit 52,000 tokens per second per GPU, versus 19,400 on the earlier stack. A 1.2T-parameter model was benchmarked on 512 GPUs.

Key Points
Experts scaled from 8 to 128 at about 3.2B active parameters, taking total size from 4.6B to 47B with under 5% throughput loss
47B MoE: 52,000 tokens/s/GPU against 19,400 before, about 2.7x
1.2T-parameter model benchmarked on 512 GPUs at 858 TFLOP/s per GPU
MXFP8 raised throughput about 21% over BF16 and cut peak memory from 103 GiB to 95 GiB
A 2.38T-parameter run with DeepEP v2 was a short capacity test, not full training
Why It Matters
Open MoE training stacks are what let academic labs follow the frontier. Code and a technical report are public, so teams can reproduce the numbers.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,556 builders reading daily.