Commerce Agent Performance Falls Short of Promise in New Alibaba Test

A recent benchmark by Alibaba.com has highlighted both the potential and limitations of AI agents in handling real commerce tasks. The CommerceAgentBench toolkit, released on GitHub , tested 13 AI model families across 107 practical scenarios with a strict pass/fail grading system – only correct job completion earned credit.

The results showed that even top performers struggled to consistently handle complex transactions. Claude Opus 5 achieved the highest score of 61.7%, which Alibaba President Kuo Zhang described as “high enough to be useful and low enough to be a warning” in a recent Fortune commentary.

The test covered critical commerce functions including procurement, logistics, product listing, fulfillment, and after-sales service. Agents had to process unstructured data like emails, detect fraud, calculate landed costs, and manage shipping across multiple carriers – all based on real tasks from Alibaba’s 10 million active small business users.

Where AI Agents Fail Most Often

The benchmark revealed several common failure points:

  • Multi-step processes: Tasks requiring agents to track information across numerous steps consistently tripped up models
  • Complex calculations: Landed cost analysis failed when multiple variables shifted simultaneously
  • Discrepancies in data: Agents struggled with after-sales disputes where documents contained conflicting information
  • Long workflows: Payment anomalies buried deep within lengthy email threads often went undetected

The top-scoring model, Claude Opus 5, averaged 63 tool calls and took about 10 minutes per task – demonstrating the computational intensity of even seemingly simple transactions.

These findings echo results from other AI benchmarks. A similar test by Mercor across white-collar domains like consulting and law found that top models correctly completed less than a quarter of tasks, with information tracking being the biggest challenge.

Task Specialization May Be Key

The Alibaba benchmark also underscored the importance of task-level routing – directing different transactions to AI models best suited for specific functions rather than relying on one general-purpose agent. This approach aligns with how human commerce teams operate, where specialists handle particular areas like logistics or customer service.

As AI adoption in commerce accelerates, these insights will be crucial for businesses seeking to deploy agents effectively while managing risks and ensuring accuracy.