
ByteDance Seed and Tsinghua AIR have released CUDA Agent, an agentic reinforcement learning system that trains a large language model to write GPU kernels that beat a compiler. The gap it targets is narrow but stubborn: frontier models already produce correct CUDA, they just produce slow CUDA. On KernelBench, the base model Seed1.6 passes 74.0% of tasks yet outruns torch.compile on only 27.2% of them, at a 0.69× geometric-mean speedup which means its kernels are, on average, slower than what th
Will ByteDance's CUDA Agent paper appear on the Hugging Face CUDA-Agent model page by August 25?
This prediction was voided because the outcome could not be settled from the source.
CUDA Agent is a reinforcement learning system that trains a large language model to write GPU code (CUDA kernels) that runs faster than compiler-generated code. Frontier AI models already produce correct GPU code but often generate slower versions, achieving only a 0.69× speedup compared to standard compilers. By placing the model in a real development environment with profiling tools and correctness checks, then training it with reinforcement learning for 150 steps, CUDA Agent improved performance to a 2.11× speedup and a 96.8% rate of outperforming the standard compiler on benchmark tasks. This matters for AI infrastructure, GPU computing, and any application where GPU kernel performance directly affects latency and cost, though the trained agent itself is not publicly released and full replication requires substantial computational resources.

Making enterprise reporting tools openly available could democratize access to automation that typically requires expensive proprietary software licenses.

AWS Strands Labs releases Strands Decider 2B, an open source decision model. It does not generate text. It reads a state and typed questions, then returns a choice, a yes/no probability, or a score with a calibrated confidence. The model has 1.9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac. Is it deployable? Yes, for local and self-hosted use. Weights are on Hugging Face under Apache-2.0, and pip install strands-decider gives a CLI and an HTTP server.

In this tutorial, we implement Kauldron, the JAX training library from Google Research that describes itself as optimized for research velocity and modularity, and we take those two words literally by testing what they actually buy us. We install it, then spend the first half of the notebook on the three mechanisms that make Kauldron different from a stack of Flax and Optax: konfig, which turns an experiment into a tree of plain dictionaries that round-trip through JSON; kontext, which wires co
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven