MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models
Abstract
In this work, we present the MaLA (Massively multilingual Language Adaptation) suite, a comprehensive collection of open-access resources designed to advance language technology across 546 languages. The primary contribution is the MaLA corpus, a 74-billion-token dataset specifically engineered for the continual pre-training of large language models (LLMs). To ensure high utility and robustness for various tasks, we introduce a diverse data mix that balances underrepresented languages with curated high-resource data. This mix includes: (1) structured scientific literature; (2) literary archives; (3) multilingual instruction sets; and (4) programming data. We provide a rigorous documentation of the corpus provenance, including our sampling strategies—upsampling low-resource languages and downsampling high-resource ones—to mitigate catastrophic forgetting while expanding language capacity. Leveraging this resource, we release EMMA-500, a Llama 2-based model optimized for cross-lingual transfer and language adaptability. Beyond the model weights, we provide the full suite of scripts, processing pipelines, and model generations to support reproducibility and further experimentation in multilingual indexing and retrieval. We release the MaLA suite, including the MaLA corpus, EMMA-500 model weights, scripts, and model generations, under open licenses to provide the community with the testbeds and tools necessary to evaluate and improve global information access.