SOURCE-LINKED INTELLIGENCE
Optimal Allocation of Embedding Dimensions under Finite-Sample Constraints
The embedding dimension of categorical predictors is usually selected through heuristic tuning, although it directly affects model complexity, approximation quality, and finite-sample generalization. This paper formulates embedding dimension selection as a constrained allocation problem. The main contribution is to show that embedding capacity can be allocated across heterogeneous categorical predictors according to an explicit approximation-estimation tradeoff. We characterize approximation error through the singular-value tail of the latent category representation, while estimation error inc
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-25T14:18:11.000Z
First collected: 2026-09-21T10:02:02.728Z. This is not the publication date.