SOURCE-LINKED INTELLIGENCE
GenEval: Linking Generation and Evaluation for Reliable NLG Assessment
GenEval: Linking Generation and Evaluation for Reliable NLG Assessment "Large Language Models (LLMs) are increasingly leveraged as evaluators of machine-generated text, a paradigm known as ""LLMs-as-judges."" While this approach offers flexibility and typically strong performance, its reliability remains inconsistent and poorly understood. Strong generative performance does not guarantee reliable evaluation, and the mechanisms linking these two capabilities remain opaque. Without systematic validation, current evaluation practices risk being blind to misleading and factually incorrect content and misrepresenting system capabilities. GenEval addresses this challenge by investigating the fundamental relationship between generation and evaluation in LLMs, developing novel representation-based metrics, and predicting LLMs' evaluation reliability across tasks and models. We will analyze LLMs at a mechanistic level, identifying circuits and representations that u
Read original source ↗ Open in workspace
- recordType
- award
- status
- SIGNED
- region
- EU
- value
- 194074.56
- unit
- EUR
Evidence & attribution
European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.
License: CORDIS reuse policy
First collected: 2026-09-20T05:31:32.981Z. This is not the publication date.