Automation Bench Verified
AutomationBench tests AI agents on business workflows in sales, marketing, human resources, operations, support and finance. It stands out among contemporary benchmarks due to its scale and uncompromising focus on realism across many business applications, which is partly why it features on many high profile model cards such as Gemini 4 Argon, Claude Opus 5.5, GPT-6 Astra, Claude Fable 5.1, and GPT-6.1 Sol
Because we admire the hard work of Daniel Shepard, Robin Salimans at Zapier we singled AutomationBench as the target of our latest audit. Our aim as always in these reviews is to highlight excellent work in the field and encourage attention to detail in public benchmarks. Importantly, we were only able to audit the 600 public tasks in this case.
The Audit
Not having access to the private set we conducted our audit on the 600 public tasks. We took the following steps:
- We conducted an agent-assisted adversarial evaluation of all tasks in which we designed adversarial submissions. These are realistic submissions intended to make the verifier produce the wrong outcome. This process flagged 323 tasks as possibly having dubious verifiers.
- We manually reviewed the flagged tasks, rejecting 117 as false positives, and accepting 206 as serious issues.
- We then applied fixes for all 206 serious issues.
In the findings section we provide a clear and detailed version of our reasoning for 100 findings so that researchers can judge for themselves whether our findings are valid. For more details on our process you can read more here.
Results
To measure whether our fixes had any meaningful impact on the tasks we audited we performed two experiments, both using the open source frontier model Kimi K3.
Firstly we identified two sets: tasks where all our fixes made the task easier and tasks where all our fixes made the task harder. In most cases our fixes removed unfair strictness and also unfair laxness in the test, so these 2 sets form a minority of all fixed tasks. Across all 206 tasks, we ran 3 Kimi K3 traces on the original pre-fixed tasks and 3 on the post-fixed tasks (a total of 1,236 trials) and recorded the average score in each case. The graph shows the two subsets: 16 tasks made less strict and 57 made more strict. In both cases the average score went in the direction of the fix.
Secondly, to test the remaining tasks we took the submissions from the 3 post-fixed runs and the 3 pre-fixed runs and judged them using the opposite verifier to see if the verdicts changed. If no verdicts change it indicates that our fixes only apply to niche cases, and maybe are not important, while if many verdicts change it indicates that our fixes have a significant impact on score. We found that 344/1,235 verdicts changed due to our fixes, or 27.9%.
Examples
High level analysis can only take us so far however. We encourage researchers to view some of the 100 findings we present here to gain a qualitative understanding of our audit. 6 indicative examples are also presented below for ease of viewing.
Conclusion
In auditing the public tasks released by AutomationBench we hope to encourage attention to detail among researchers and benchmark auditors. We believe we have somewhat substantiated the idea there there are a significant number of issues in this task set and that it would benefit from implementing our fixes. Ultimately our investigation indicates that AutomationBench is a useful and powerful tool for guiding the inception of AGI and if given the opportunity we would love to collaborate with Zapier in performing a similar audit of the private set as well.
Scope: The 206 findings include both scored and unscored public tasks.
