
Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments
Google has released an updated AI weather model that produces forecasts based on real-time satellite observations rather than relying solely on traditional physics-based simulations. The new model generates predictions every hour at five-kilometer resolution, compared to the previous version's six-hour forecasts at 25-kilometer resolution, which is particularly useful for predicting fast-moving rain and snow systems. AI weather models recognize patterns in historical data to make faster predictions than traditional supercomputers, though they are still expected to work alongside physics-based models used by weather agencies. The improved model also produces forecasts for renewable energy generation, such as wind speed predictions at turbine height, and is now integrated into Google's search, maps, and other products.

Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from

Time to first token (TTFT) is the metric teams use to pick an inference API for voice. It is also the metric that misleads them. TTFT marks when generation starts; a text-to-speech model cannot speak until a full clause arrives. Between those two points sits the difference between an agent that feels conversational and one that gets interrupted. This piece benchmarks every layer of the voice stack including LLM, speech-to-text, text-to-speech, and speech-to-speech. Why TTFT Is the Right Entr
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven