Wireva

Alibaba President: AI agents can talk, but can they actually do the work?

's president argues that the industry's focus on model intelligence misses the point. Real commercial work requires agents that can execute tasks correctly, not just converse. A new open-source benchmark, CommerceAgentBench, grades outcomes and reveals that even the strongest frontier models fail nearly 40% of real-world e-commerce tasks.

This item was produced with AI assistance under the editorial responsibility of Haydamax OÜ.

Monitoring item. The full text is not distributed. Extract and source below.

's president argues that the industry's focus on model intelligence misses the point. Real commercial work requires agents that can execute tasks correctly, not just converse. A new open-source benchmark, CommerceAgentBench, grades outcomes and reveals that even the strongest fr…