Researchers are testing whether reinforcement learning techniques designed for AI agents can teach models to master creative tasks beyond their original training data.
A coding model was trained to paint watercolors by writing JavaScript code that uses drawing tools to create images. The training used reinforcement learning guided by aesthetic preferences rather than a single correct answer, combining scores from a compiled code checker, code length measure, a vision model comparing paintings to hand-rated reference images, and a preference model trained on human image choices. The approach treats the hand-curated pool of reference images as the definition of what is beautiful, making the work about building that set rather than tuning technical parameters. All components of this system, including the training scripts, environment, models, and datasets, were published openly on Hugging Face.

While restaurant owners might look to generative AI as a shortcut to sprucing up their menu, customers can viscerally sense that something is wrong with the food.
.gif&s=kx7H9a1S3ZIMP7WHNPT2XBlIq6uwDsVvpMZbLf-dnSA)
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.

Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now released that scaffolding as code. This week, they published anthropics/commerce-agents, a reference blueprint containing a shopping agent and a merchant agent, along with four runnable verticals: retail, travel, telecom and entertainment. It ships alongside two write-ups: a product announcement and an engineering de
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven