SOURCE-LINKED INTELLIGENCE
Assessing Training Data Filtering Strategies for the Reduction of Harms in AI
Assessing Training Data Filtering Strategies for the Reduction of Harms in AI Large Language Models (LLMs) learn everything they know of the World from what they find in training datasets: if datasets include harmful content, it is more likely that they learn how to produce discriminating outputs. Therefore, reducing the presence of harmful contents in the input training dataset is a crucial step to develop safer and fairer technologies. However, the effectiveness of existing data filtering strategies for harm reduction is still an understudied topic in NLP research. DISHARM's main research objective is to develop the first framework for the systematic evaluation of data filtering strategies with the goal of reducing harmful contents in training datasets. The project foresees the implementation of the first open leaderboard of data filtering strategies, which will enable a comparison of their effects in mitigating or exacerbating harms against different vulnerabl
Read original source ↗ Open in workspace
- recordType
- award
- status
- SIGNED
- region
- EU
- value
- 263393.28
- unit
- EUR
Evidence & attribution
European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.
License: CORDIS reuse policy
First collected: 2026-09-20T04:21:15.460Z. This is not the publication date.