Lang-Prune: Unlocking Fair and Powerful Pruning for Multilingual Large Language Models
Juhao Liang ⋅ Shiqi Zhang ⋅ Min Zhang ⋅ Hao Yang ⋅ Benyou Wang
Abstract
Multilingual large language models (LLMs) are essential for cross-lingual applications, yet pruning them with mixed-language calibration induces *cross-lingual interference*, disproportionately degrading low-resource and script-diverse languages. We introduce ***Lang-Prune***, a drop-in extension to structured pruning that computes importance per language on small calibration sets and aggregates via a Max operator---preserving any structure critical to at least one language. On *aya-expanse-8b* across nine typologically diverse languages, Lang-Prune reduces average perplexity by 62\% at 70\% sparsity relative to mixed-data pruning, and outperforms monolingual pruning on average under the same calibration budget. Lang-Prune generalizes across model families and scales: its advantage amplifies with model size, reaching $4.1\times$ perplexity reduction at *Qwen3-14B* and exhibiting strong zero-shot transfer to out-of-distribution languages. Post-training experiments further show preserved adaptation capacity, with pruned models reaching 59.6\% vs. 50.6\% accuracy on Belebele. Lang-Prune modifies only the importance estimation and aggregation step, preserving LLM-Pruner's coupled-structure mechanics and requiring no additional parameters or retraining.
Successful Page Load