AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings

arXiv · AI, language, vision and robotics · article · Sep 12, 2026 · UTC

Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each respo

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T16:41:15.630Z. This is not the publication date.