The strongest frontier model we tested successfully completed 61. 7% of the tasks — high enough to be useful and low enough to be a warning.
Source: [Fortune](https://fortune.com/2026/09/09/alibaba-president-kuo-zhang-ai-agents-can-talk-but-not-work/)