
Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models. It offers serverless endpoints and reserved capacity on Prime’s own GPUs across multiple datacenters. Before public release, it processed nearly a trillion tokens per day internally. That traffic came from RL rollouts, synthetic data generation, evaluations and long-running coding agents. What is Prime Inference? Prime Inference is the serving layer of Prime Intellect’s open train
Prime Intellect has launched Prime Inference, a platform for running open-source AI models that offers both serverless endpoints for variable demand and reserved capacity for sustained workloads. The platform processes nearly a trillion tokens per day internally from training-related tasks like synthetic data generation and coding agents. It uses a technical stack combining multiple open-source components to optimize performance, achieving features like automatic failover across datacenters and near-zero tool-call error rates for AI agents. The service is OpenAI-compatible and offers unified billing, though full pricing details are not yet published.

Millions have downloaded Meta’s AI agent Muse. But getting it to do your bidding comes with privacy costs.

We created a list of the most notable AI agents that can live in your text messages, from general assistants to agents designed for families, travel, and work.

TL;DR: Meta’s Muse, OpenAI’s Dots and Uber’s driver assistant share one bet: the agent speaks first. That moves the hard problem from what to answer to when to interrupt, on which channel, and with what offer. Classic ML and new decision models can solve it. Three launches, one pattern Meta Muse (Sept 8). A personal agent that books, emails and keeps working with the app closed. It remembers details, makes unprompted suggestions and checks in for approval, in its own ap
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven