benchmarkscritiquebearishComputer-use agent benchmark scores are commonly produced by brittle scripted oracles that can produce unreliable resultsArtificial Intelligence02 Aug 2026http://arxiv.org/abs/2607.28367v1