SOURCE-LINKED INTELLIGENCE
Just add noise: Debiasing tree-based variable importance in mixed data
Variable importance scores from tree-based methods such as random forests favor continuous predictors over categorical ones. We present a theoretical analysis of this bias and propose a simple remedy: add a small amount of noise to each categorical predictor. The correction is demonstrated on a variety of simulated and real-world datasets and combined with integrated path stability selection to perform variable selection with false discovery control for mixed data.
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-12T17:53:28.000Z
First collected: 2026-09-20T12:41:04.663Z. This is not the publication date.