
Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
Google has introduced Gemini 3.5 Transcribe, a speech-to-text model that converts spoken audio into written text while handling background noise, specialized vocabulary, and multiple languages automatically. Unlike traditional speech recognition, the model cleans up natural speech patterns by removing filler words and self-corrections, and can understand context to improve accuracy. The model is available to developers through APIs for both real-time voice applications and recorded audio processing, and is being integrated into Google products including the Gemini app, Android, and Chrome. According to the company's measurements, the model achieves lower error rates than its predecessor and performs better across noisy environments and over 85 languages.

While restaurant owners might look to generative AI as a shortcut to sprucing up their menu, customers can viscerally sense that something is wrong with the food.
.gif&s=kx7H9a1S3ZIMP7WHNPT2XBlIq6uwDsVvpMZbLf-dnSA)
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.

Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now released that scaffolding as code. This week, they published anthropics/commerce-agents, a reference blueprint containing a shopping agent and a merchant agent, along with four runnable verticals: retail, travel, telecom and entertainment. It ships alongside two write-ups: a product announcement and an engineering de
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven