benchmarkscritiquebearishSWE-bench-like benchmarks suffer from systematic misalignment due to the complexity of PR-Issue pairing in large repositoriesArtificial Intelligence02 Aug 2026http://arxiv.org/abs/2607.28587v1