OSWorld 2.0 Verified
We decided to turn our internal quality control pipeline on one of our favorite benchmarks: OSWorld 2.0. We follow the excellent work in Introducing SWE-bench Verified by using direct human assessment to attempt to identify any issues in SWE-bench. Our goal was to detect the following classes of issues:
- Major: There exist broad range of plausible cases where the task does not measure what it is supposed to measure.
- Minor: Improvements are possible, but the task still mostly measures the intended behavior.
We reviewed all 108 tasks and retained 43 findings: 18 major and 25 minor. Each case links the submitted files and available exhibits so readers can inspect the result.
For instance from task 092 requires a spinning 3D logo, but a correctly styled 3D logo that never rotates still gets full credit.
Golden submission
Adversarial example
Both these submissions score a full 1.0. We consider this a major defect because an agent could plausibly forget to create the rotation effect altogether and still receive full marks. In this case the task does not test what it is supposed to.
Another major issues appears in task 61 where the agent is supposed to apply a color grade to a snowy image. The verifier is very forgiving however, and other transforms which generate very incorrect images, such as a per-channel affine color match followed by a blur also score 1.0.

An example of a minor issue can be found in task 50 where an agent can receive full credit for an edited podcast which accidentally mutes one of all of the tracks. This is not ideal, but since agents are highly unlikely to do this in practice we don't think this will seriously effect the task's usefulness.
Golden submission
Adversarial example
Our hope is that researchers using this benchmark can use our findings to get even more out of the excellent tool that is OSWorld 2.0.
NOTE: we are auditing release v2026.08.08 which is the most recent release at time of writing (10/09/2026), though we believe the OSWorld authors are working on a new release.
Task ID lists
Tasks with major issues (18)
007, 017, 020, 037, 043, 045, 046, 051, 056, 061, 087, 088, 089, 091, 092, 093, 095, 102
Tasks with minor issues (25)
001, 002, 005, 008, 018, 023, 026, 038, 042, 044, 047, 049, 050, 060, 062, 063, 070, 071, 084, 085, 096, 098, 106, 107, 108
Tasks with issues (43)
001, 002, 005, 007, 008, 017, 018, 020, 023, 026, 037, 038, 042, 043, 044, 045, 046, 047, 049, 050, 051, 056, 060, 061, 062, 063, 070, 071, 084, 085, 087, 088, 089, 091, 092, 093, 095, 096, 098, 102, 106, 107, 108
References
- Yuan, M., Zhou, Z., Xiong, X., et al. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. arXiv:2606.29537, 2026.
- OpenAI. Introducing SWE-bench Verified. 2024.
Findings.
All findings.
All — findings upheld after individual adjudication.
Pending classification
How we tested the evaluators.
Every finding published here has passed through multiple sets of human eyes.
Coming soon.
Case not found
That evidence record is not in the projection.
Return to Findings to browse the canonical package cases.
Open Findings