
Lightweight model versions enable vision and reasoning tasks on edge devices with limited compute, extending AI beyond cloud-dependent systems.
A world model is a type of artificial intelligence that learns how environments change over time by understanding objects, motion, spatial relationships, and the effects of actions. Cosmos 3 Edge is a 4-billion-parameter open world model designed to help robots and vision AI agents understand their surroundings and generate actions on edge devices, which are memory-constrained computers used in factories, warehouses, hospitals, and similar settings. The model uses two transformer towers with shared attention layers, allowing it to reason about current scenes, predict possible futures, and connect those predictions to physical actions through a unified representation. Developers can download the model and adapt it for their own applications using provided training scripts and post-trained checkpoints.

In this tutorial, we explore FAIRChem v2 and the UMA universal machine-learning interatomic potential as a unified framework for atomistic simulation across molecular chemistry, catalysis, and inorganic materials. We configure an environment, authenticate with Hugging Face to access the gated UMA model weights, and initialize task-specific calculators for the omol, oc20, and omat domains. We then apply the same pretrained potential to a broad set of computational chemistry workflows, including

Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer trained on 18 years of computer demonstration video. On an internal computer use benchmark, Induction Labs reports that Photon-1

The KwaiKAT Team at Kuaishou has introduced the KAT-Coder-V2.5. It is a coding model trained to operate inside real, executable repositories rather than emit single-turn code. The served model is available through StreamLake. An open-weight variant, KAT-Coder-V2.5-Dev, was released separately on Hugging Face under Apache-2.0. AutoBuilder: environments that actually run the intended tests The research frames a verifiable task as a triplet. It needs a precise task description, an executable
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven