benchmarksfactbullishGemini 3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement for grading command-line examinations among frontier LLMs (GPT, Claude Opus, Gemini, GLM)Computation and Language27 Jul 2026http://arxiv.org/abs/2607.02432v1