benchmarksopinionbearishGrammatical competence in morphologically rich languages remains under-measured in the rapidly scaling open-weight LLM ecosystemMachine Learning02 Aug 2026http://arxiv.org/abs/2607.28274v1