source: Hugging Face Blog: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

level: technical

allenai released olmo-core 3, an open training framework for large mixture-of-experts language models. it replaces the earlier fully sharded data parallel approach with distributed data parallel, keeping experts resident on gpus and routing data to them. this avoids repeated weight gathering and supports scaling to over one trillion total parameters. the release includes code, a technical report, and an interactive demo.

in benchmarks on eight nvidia b300 gpus, a 47-billion-parameter moe processed 52,000 tokens per second per gpu with the new stack, versus 19,400 with the previous implementation, a 2.7x improvement. scaling expert count from 8 to 128 while keeping four active per token raised total parameters from 4.6b to 47b with under 5% throughput loss. mxfp8 precision added 21% throughput over bf16, reducing peak memory from 103 to 95 gib.

the framework combines expert, pipeline, and data parallelism with a distributed optimizer, plus techniques like rowwise expert parallelism and grouped gemm. tests reached 1.2 trillion parameters across 512 gpus and a short-capacity run at 2.38 trillion. the report documents negative results, such as token gerrymandering in routing balance scores and cases where overlapping communication slowed training. this open stack lets smaller labs train large moes efficiently.

why it matters: open, efficient moe training infrastructure lowers compute costs and energy use, making trillion-parameter model development accessible to academic and smaller labs.


source: Hugging Face Blog: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs