PCA-guided Activation Scaling for Monotonic Bidirectional Control over LLM Sycophancy
Abstract
Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirely risks over-correction against valid opinions. Effective control must therefore both reduce and increase sycophancy with predictable and gradual effect. Yet, existing methods fail to ensure a bidirectional and monotonic relationship between steering strength and behavioral outcome across models and datasets. We introduce PCA-guided Activation Scaling (PAS), an activation steering framework that decomposes residual stream activations into a PCA-identified sycophancy-honesty subspace and an orthogonal residual, then applies distinct scaling exponents to achieve monotonically bidirectional control. Across three LLMs and three datasets, PAS achieves strong monotonicity (Spearman rho = 0.92) and an average shift of 15.4% per direction, compared with rho =-0.05 and 8.7% for the baseline. Ablation studies confirm that the decomposition, asymmetric exponents, and layer selection are each essential for maintaining monotonic control. The data and code is available at \url{https://anonymous.4open.science/r/PCS-D0D8/}.