
Open-source alternatives to proprietary model training reduce barriers for researchers and smaller organizations building competitive AI systems.
Olmo-core 3 is an open framework for training large mixture-of-experts language models, which are AI systems that use specialized component parts selectively rather than activating all parameters for every input. Training such large models typically requires substantial computing resources, but mixture-of-experts architectures promise greater efficiency by using only relevant experts per input while still containing many total parameters. The new system addresses a key challenge: as these models grow larger, the overhead of directing data to the right experts and coordinating across GPU clusters can eliminate much of the efficiency gains. Olmo-core 3 combines several techniques including expert parallelism, pipeline parallelism, and a distributed optimizer to allow mixture-of-experts models to scale into the trillion-parameter range while maintaining computational efficiency, and it has been benchmarked to achieve significant throughput improvements compared to earlier approaches.

Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, not chatbots. Each reads an input state and a schema of typed questions. It returns a probability for every allowed answer, with no free-form text. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI’s Jev API. Is it deployable? Yes, Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting. What a Decision Model Do

In this tutorial, we implement Kauldron, the JAX training library from Google Research that describes itself as optimized for research velocity and modularity, and we take those two words literally by testing what they actually buy us. We install it, then spend the first half of the notebook on the three mechanisms that make Kauldron different from a stack of Flax and Optax: konfig, which turns an experiment into a tree of plain dictionaries that round-trip through JSON; kontext, which wires co

Cohere has released Embed 5, a new embedding model family. It targets enterprise search, RAG, and agentic retrieval. The model family ships in 2 tiers. Embed 5 Pro targets maximum retrieval quality. Embed 5 Fast targets latency and cost on the live query path. Both accept text, images, and fused text plus image inputs. Both cover 100+ languages and read up to 128K tokens. The key design choice: Pro and Fast share 1 embedding space. You can index with one and query with the other. Is it deplo
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven