Papers
arxiv:2606.00079

BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization

Published on Sep 27
Authors:
,
,
,
,
,

Abstract

BitsMoE is a spectral-energy-guided bit-allocation framework that reduces memory requirements for Mixture-of-Experts language models through SVD decomposition and mixed-precision quantization while maintaining accuracy in ultra-low-bit regimes.

Mixture-of-Experts (MoE) large language models incur substantial memory costs due to their large expert parameter counts. Mixed-precision quantization reduces these costs by allocating different bit-widths to experts or linear blocks according to their importance. However, assigning a single precision within each expert or linear block overlooks its internal structural heterogeneity. This limitation motivates two key questions: (1) how to define a fine-grained unit for quantization within a linear transformation; and (2) how to characterize the quantization cost of each unit under actual activation patterns and different bit-widths. To address these two questions, we propose BitsMoE, a cost-aware mixed-precision quantization framework built on two complementary techniques: (1) Shared-basis Spectral Decomposition (SSD) separates expert weights into a shared basis and expert-specific spectral components, defining structural quantization units while exploiting cross-expert redundancy. (2) Factorized Quantization Cost Modeling (FQCM) estimates component-wise costs from output reconstruction loss by combining intrinsic spectral importance, activation-dependent importance, and bit-width-dependent distortion. Using these component-wise costs, we formulate bit allocation as an integer linear program (ILP) that minimizes total modeled quantization cost under a fixed memory budget. On Qwen3-30B-A3B at 2-bit, BitsMoE achieves 64.29% average accuracy over seven downstream tasks, outperforming the evaluated MoE-specific methods, including those using ILP-based bit allocation, and exceeding GEMQ by 2.80 percentage points. Under the same setting, it achieves a 16.47times end-to-end offline quantization speedup over GEMQ. It also achieves up to 6.46times the decode throughput of GPTQ.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.00079
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 4

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.00079 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.00079 in a Space README.md to link it from this page.

Collections including this paper 1