
Smaller, more efficient model versions could make advanced AI more accessible to researchers and developers with limited computing resources.
Quantization-Aware Distillation (QAD) is a technique that compresses AI models into smaller 4-bit versions by training a smaller "student" model using knowledge from a larger "teacher" model, rather than shrinking an already-trained model. This matters because the resulting compressed models run on edge devices like phones and laptops with the same speed and memory efficiency as standard compression methods, but recover 97 percent of the accuracy lost in the compression process. The LFM2.5 models in four sizes have been released as QAD checkpoints that can be used with standard AI inference software.

What are Decision AI Models? Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly. The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking. Within 3 weeks, Fastino Labs shipped 2 rival

Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience. What Microsoft Shipped MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-

Making enterprise reporting tools openly available could democratize access to automation that typically requires expensive proprietary software licenses.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven