MoANT: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning
Abstract
Large language models (LLMs) encounter significant adaptation challenges in diverse multitask finetuning. Mixture-of-experts (MoE) provides a promising solution with a dynamic architecture, enabling effective task decoupling. However, scaling up the number of MoE experts incurs substantial parameter and computational overheads and suffers from limited performance gain due to naive routing mechanisms. In this paper, we design a novel framework, mixTure-of-Rank-onE-eXpert (T-REX), which leverages the combination of ultra-low rank experts to construct LoRA weights on pretrained LLMs. The rank-1 experts enable a mix-and-match mechanism to quadratically expand the vector subspace of experts with linear parameter overheads, achieving approximate error reduction with optimal efficiency. In addition, T-REX provides implicit guidance to the router, leveraging the inherent semantic clustering of training embeddings as prior knowledge, which enables optimized feature allocation across experts for smoother convergence. T-REX demonstrates superior efficiency and generalization across diverse in- and out-of-distribution tasks. Compared to LoRA, it achieves up to a 1.78% gain in mean accuracy while using approximately 40% fewer trainable parameters across 14 public datasets and 8 LLM backbones.