
Liquid AI has announced LFM2.5-VL-3B-DSpark, an experimental speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The drafter adds about 280M parameters and speeds up decoding without changing the model’s output. Liquid AI team reports up to 3.13x faster decoding on Apple silicon and up to 2.66x on an NVIDIA H100. Is it deployable? Yes, Weights are live on Hugging Face in Safetensors and GGUF, with day-one support in SGLang, MLX-VLM, and llama.cpp. Liquid AI tea
Speculative decoding is a technique that speeds up vision-language models by using a smaller "drafter" model to propose multiple tokens ahead, which a larger model then verifies in one pass instead of generating tokens one at a time. A new drafter model has been released that adds about 280 million parameters to an existing vision-language model and achieves up to 3.13 times faster decoding on Apple silicon and up to 2.66 times faster on an NVIDIA H100. The drafter works by reading hidden states from the target model's layers and predicting the next several tokens, with the key insight being that the technique works regardless of whether the input is text or image data since both are represented as tensors in the hidden layers. The speedup is most noticeable during the decoding phase, while image encoding and initial processing remain at the same speed, which limits the total end-to-end acceleration especially on edge devices where these other steps take up a larger share of the total time.

Google Research has announced a next-generation Federated Learning (FL) system built on Trusted Execution Environments (TEEs). The research team claims externally verifiable central differential privacy (DP) guarantees for FL for the first time. What Problem Does TEE-Based Federated Learning Solve? Google introduced Federated Learning FL in 2017. It powers next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages, and Smart Text Selection in Android. Earli

Aleph Alpha has released Kolibri, an open-weight Mixture-of-Experts (MoE) language model built for German and English. Kolibri has 78.1B total parameters but activates only 3.46B, or 4.4%, per token. It accepts up to 1,048,576 tokens of context, lets users set reasoning effort per request, and ships under the Apache 2.0 license on Hugging Face. The target is sovereign deployment in regulated sectors such as public administration, industry and aerospace. Is it deployable? Yes. The FP8 checkpo

What are Decision AI Models? Decision AI models are a new class of model that returns a decision, not a paragraph. You send text with typed questions. The model returns choices, scores or yes/no probabilities your code can branch on directly. The category went mainstream when TypeSafe AI launched Jev after 2 years in stealth. TypeSafe calls it a ‘System One model,’ after Daniel Kahneman’s fast, intuitive System 1 thinking. Within 3 weeks, Fastino Labs shipped 2 rival
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven