WORLDTECH NEWS Global technology intelligence.Contact
← Back to WORLDTECH

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Rows of cooling units on a data-center roof at sunsetAI illustration
WORLDTECH illustration · AI-generated (Canva)

What happened

Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs.

MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU (a chip built for many calculations at once, used for graphics and AI) memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs.

As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input. Olmo-core 3 is built to close that gap. Total parameter capacity grew from 4.6B to 47B, while training throughput (how much work a system gets through in a given time) fell by less than 5%.

Sources & evidence