AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC

Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate th

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:02:05.452Z. This is not the publication date.