agentscritiquebearishLarge language models can synthesize executable Rust tests but their outputs often violate API preconditions, remain shallow, or reduce concurrency to accidental sequential tracesArtificial Intelligence29 Jul 2026http://arxiv.org/abs/2607.21530v1