AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Sustainable Training of Code Language Models through Data Refinement

CORDIS · observation · Publication date unknown

Sustainable Training of Code Language Models through Data Refinement "Large language models (LLMs) have gained widespread attention and user adoption. These models, when trained on source code from platforms like GitHub, acquire a deep understanding of both the semantic and syntactic structures of code (i.e., code language models or CLMs). This understanding has paved the way for significant advancements in software engineering, offering developers valuable assistance in labor-intensive tasks like bug fixing and code writing. While CLMs offer tremendous assistance in software engineering tasks, their massive data requirements result in substantial energy consumption and CO2 emissions. This proposal challenges the conventional wisdom that ""more data is better"" and instead advocates for a refined approach to data in the training of CLMs. We propose that by intentionally decreasing training data volume while simultaneously enhancing data quality through da

Read original source ↗ Open in workspace

recordType
award
status
SIGNED
region
EU
value
210911.04
unit
EUR

Evidence & attribution

European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.

License: CORDIS reuse policy

First collected: 2026-09-20T02:21:08.944Z. This is not the publication date.