AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Assessing Training Data Filtering Strategies for the Reduction of Harms in AI

CORDIS · observation · Publication date unknown

Assessing Training Data Filtering Strategies for the Reduction of Harms in AI Large Language Models (LLMs) learn everything they know of the World from what they find in training datasets: if datasets include harmful content, it is more likely that they learn how to produce discriminating outputs. Therefore, reducing the presence of harmful contents in the input training dataset is a crucial step to develop safer and fairer technologies. However, the effectiveness of existing data filtering strategies for harm reduction is still an understudied topic in NLP research. DISHARM's main research objective is to develop the first framework for the systematic evaluation of data filtering strategies with the goal of reducing harmful contents in training datasets. The project foresees the implementation of the first open leaderboard of data filtering strategies, which will enable a comparison of their effects in mitigating or exacerbating harms against different vulnerabl

Read original source ↗ Open in workspace

recordType
award
status
SIGNED
region
EU
value
263393.28
unit
EUR

Evidence & attribution

European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.

License: CORDIS reuse policy

First collected: 2026-09-20T04:21:15.460Z. This is not the publication date.