safety
fact
bearish
JUDGESTEALER can achieve up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation respectively when extracting LLM judge capabilities
Extensive experiments on state-of-the-art LLM-as-a-judge and reward models show that JUDGESTEALER consistently outperforms existing extraction baselines, achieving up to 73.3%, 87.0%, and 71.6% accuracy for pointwise, pairwise, and listwise evaluation, respectively.
Computation and Language29 Aug 2026