
Faster vision-language models could enable real-time analysis of images and video across more devices and applications.
A vision-language model draft model has been released that speeds up inference by using speculative decoding, a technique that trades a small increase in memory for faster processing without changing output quality. The drafter adds approximately 280 million parameters, increasing the target model's size by 8.9%, while delivering decoding speedups up to 3.13x on device and up to 2.66x on GPU hardware. Speculative decoding works by having the draft model capture hidden states from the target model and propose multiple candidate tokens at once, which the target model then verifies. The speedup matters most for the decoding phase, though overall improvements are more modest when vision encoding and initial processing stages already consume significant time, a limitation explained by how different computational bottlenecks affect different parts of the inference pipeline.

Google Research has announced a next-generation Federated Learning (FL) system built on Trusted Execution Environments (TEEs). The research team claims externally verifiable central differential privacy (DP) guarantees for FL for the first time. What Problem Does TEE-Based Federated Learning Solve? Google introduced Federated Learning FL in 2017. It powers next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages, and Smart Text Selection in Android. Earli

Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace. Is it deployable? Yes. The FP8 checkpo

What are Decision AI Models? Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly. The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking. Within 3 weeks, Fastino Labs shipped 2 rival
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven