AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

In this paper, we present a Mixture-of-Experts (MoE) quantization method based on activation entropy. Although quantization reduces memory and computational costs, it can substantially degrade performance. In particular, performance decline is pronounced in quantized MoE models, where individual experts have a small number of parameters that are sensitive to low-bit representation. Considering that MoE operates as an ensemble model with collaborative contributions from routed experts, a significant performance decline of a particular expert due to quantization can harm model performance. There

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:20:57.646Z. This is not the publication date.