
Quantitative research agents that write their own experiments can corrupt the evidence they later learn from. A leaky feature that scores well gets stored as a successful precedent and propagated through later iterations. Prompt-level instructions and reviewer agents do not close this, because author and reviewer share the same blind spots. A team of researchers from Princeton University, Ant Group and Stanford University propose AQuA. AQuA is a pair of language-model-driven research systems th
AQuA is a framework that uses AI agents to automatically discover trading factors and develop predictive models for financial markets while preventing a common research failure: agents learning from their own experimental mistakes instead of genuine discoveries. The system works in two independent parts, one for discovering crypto trading signals and another for building stock return prediction models, with a key design principle that agents can explore freely within their designated tools but cannot change the standards by which their work is evaluated. This matters because quantitative research is vulnerable to subtle methodological errors that produce misleading backtest results, and traditional safeguards like human review fail when both the creator and reviewer share the same blind spots. The framework addresses this by freezing the evaluation criteria and data splits before any iterations begin, ensuring that what improves across iterations is the research process itself rather than the agent's ability to exploit flaws in the evaluation system.

While restaurant owners might look to generative AI as a shortcut to sprucing up their menu, customers can viscerally sense that something is wrong with the food.
.gif&s=kx7H9a1S3ZIMP7WHNPT2XBlIq6uwDsVvpMZbLf-dnSA)
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.

Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an eval suite. Anthropic has now released that scaffolding as code. This week, they published anthropics/commerce-agents, a reference blueprint containing a shopping agent and a merchant agent, along with four runnable verticals: retail, travel, telecom and entertainment. It ships alongside two write-ups: a product announcement and an engineering de
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven