
The partnership lowers barriers for companies wanting to customize generative models for their specific visual content without extensive machine learning expertise.
NVIDIA NeMo Automodel is an open-source library that enables fine-tuning of diffusion models (used for generating images and videos) at scale without converting model checkpoints or rewriting code. The collaboration between NVIDIA and Hugging Face brings production-grade distributed training capabilities to any Diffusers-format model on the Hugging Face Hub, supporting both full fine-tuning and parameter-efficient approaches like LoRA. This matters because training and fine-tuning diffusion models require memory-efficient sharding, latent caching, and configurations that scale from one GPU to hundreds, capabilities that NeMo Automodel provides through parallelism options declared via configuration rather than code rewrites. The integration works directly with pretrained models from the Hub and allows fine-tuned checkpoints to load back into standard Diffusers pipelines for inference or sharing without additional conversion steps.

In this tutorial, we explore FAIRChem v2 and the UMA universal machine-learning interatomic potential as a unified framework for atomistic simulation across molecular chemistry, catalysis, and inorganic materials. We configure an environment, authenticate with Hugging Face to access the gated UMA model weights, and initialize task-specific calculators for the omol, oc20, and omat domains. We then apply the same pretrained potential to a broad set of computational chemistry workflows, including

Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imagination models, a foundation model architecture that pretrains on raw video with no action labels at all. Their test system is Photon-1, a sparse 106B-A5B mixture-of-experts (MoE) transformer trained on 18 years of computer demonstration video. On an internal computer use benchmark, Induction Labs reports that Photon-1

The KwaiKAT Team at Kuaishou has introduced the KAT-Coder-V2.5. It is a coding model trained to operate inside real, executable repositories rather than emit single-turn code. The served model is available through StreamLake. An open-weight variant, KAT-Coder-V2.5-Dev, was released separately on Hugging Face under Apache-2.0. AutoBuilder: environments that actually run the intended tests The research frames a verifiable task as a triplet. It needs a precise task description, an executable
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven