
OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.
Will OpenAI publish a detailed cost-per-token reduction figure in a public report by September 30, 2026?
Resolves by Sep 30, 2026
OpenAI has released performance results from Jalapeño, its first custom inference chip designed to run AI models more efficiently. The chip achieved higher throughput per kilowatt and lower latency than existing commercial systems when tested on multiple model families, demonstrating advantages in speed and energy efficiency. This matters because OpenAI believes that integrating custom chips with software, models, and data centers creates a system where improvements in one area strengthen the others, ultimately delivering more useful intelligence from each unit of compute at lower cost. The company uses this integrated approach alongside partnerships with multiple hardware and cloud providers to balance capability, speed, efficiency, and cost across different AI workloads.

Runable says 60%–70% of its 1 trillion-plus token usage in the last 90 days came from paying customers.

Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.

Perplexity has released Portable Computer, a local-first build of its agentic Computer platform that runs the agent harness, orchestrator, planner, tool router and post-trained models directly on NVIDIA DGX Spark. The local model, inference engine, tool sandbox and app connectors ship as one packaged system, every task begins on the device, and work handled by local models carries no per-token charge. When a step needs the live web or frontier reasoning, the orchestrator stops and asks before s
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven