Skip to content
daily-hour-news·

🔬Ai2's Olmo-core 3 Trains MoEs 2.7x Faster Per GPU

TL;DR

Ai2 released Olmo-core 3, open training infrastructure for large mixture-of-experts models. A 47B MoE hit 52,000 tokens per second per GPU, versus 19,400 on the earlier stack.

Ai2 released Olmo-core 3, open training infrastructure for large mixture-of-experts models. A 47B MoE hit 52,000 tokens per second per GPU, versus 19,400 on the earlier stack. A 1.2T-parameter model was benchmarked on 512 GPUs.

Ai2's Olmo-core 3 Trains MoEs 2.7x Faster Per GPU — daily-hour-news

Key Points

1

Experts scaled from 8 to 128 at about 3.2B active parameters, taking total size from 4.6B to 47B with under 5% throughput loss

2

47B MoE: 52,000 tokens/s/GPU against 19,400 before, about 2.7x

3

1.2T-parameter model benchmarked on 512 GPUs at 858 TFLOP/s per GPU

4

MXFP8 raised throughput about 21% over BF16 and cut peak memory from 103 GiB to 95 GiB

5

A 2.38T-parameter run with DeepEP v2 was a short capacity test, not full training

Why It Matters

Open MoE training stacks are what let academic labs follow the frontier. Code and a technical report are public, so teams can reproduce the numbers.

Quick Facts

Ai2Olmomixture of expertsMXFP8training infrastructureopen sourceHugging Face

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,556 builders reading daily.

Also get