Sememe: Causal Modeling for Task-Aware Depth Pruning of Large Language Models
Abstract
Large Language Models (LLMs) exhibit strong generalist capabilities, yet task-specific behavior is often supported by diffuse and redundant structures. While pruning has been widely explored as a means to reduce inference costs, few existing methods translate theoretical sparsity into tangible speedups without sacrificing commensurate performance. We introduce Sememe, a task-aware depth-pruning method that identifies unnecessary transformer blocks by using a Structural Causal Model (SCM) to derive a provably unbiased estimator that jointly exploits per-example loss variance. By estimating this causal impact, Sememe surgically identifies redundancy using a small calibration set and only the forward pass, avoiding the memory overhead of gradients or activations. Experiments across natural language benchmarks show that Sememe consistently outperforms prior approaches by utilizing calibration data more effectively to produce better-fitting performance predictions. Our approach enables high-performance, task-specific deployments with reduced latency and memory demands, achieving speedups proportional to the fraction of removed layers while retaining more downstream performance than comparable training-free depth-pruning methods.