
Autonomous video analysis capabilities enable AI systems to extract insights and take actions based on visual content without human intervention.
Google DeepMind has launched a new video analysis feature that allows AI models to actively search through video content rather than passively viewing it frame-by-frame. Instead of processing every frame at a fixed rate, the agentic video understanding capability lets the model decide what to watch, at what speed, and whether to examine video frames, audio, or transcripts. This approach reduces token consumption by up to 88 percent and costs by up to 66 percent while improving accuracy by up to 7 percent, with particularly large efficiency gains on long-form video content. The feature is now available through the Gemini API and will expand to the Gemini app and YouTube's Ask YouTube feature.

AI weather models have spent three years closing the gap with physics-based forecasting, but two problems stayed open: resolution too coarse for local terrain, and initialization tied to numerical weather prediction (NWP) analysis that arrives about six hours late. WeatherNext 3, released by Google DeepMind and Google Research, attacks both. It takes a live global geostationary satellite mosaic as a direct model input, re-initializes every hour, and emits forecasts down to 0.05° (~5 km) while t

Today, OpenAI released GPT-6 Astra. The company calls it its most intelligent and aligned model, and positions it primarily as a computer-use system rather than a chat model. The pitch is that Astra operates software the way a person does, across browsers, spreadsheets, desktop applications and terminals, and finishes multi-step jobs instead of describing how to do them. Is it deployable? Partly, and not on your own hardware. Astra is a closed, hosted model with no released weights, so self
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven