Auditing Moderation Robustness Under Realistic Settings
Abstract
Large language models are widely used for content moderation (CM) on social media platforms to ensure the incoming text content follows the platform policy guidelines. These models can be vulnerable to adversarial perturbations leading to misclassification. We design an auditing framework to systematically evaluate the empirical robustness of CM models against various input perturbations that try to either bypass a malicious content or suppress a benign content without changing the original content semantics. We evaluated the robustness of multiple off-the-shelf CM models across different public datasets and studied the fairness of robustness scores across different ethnicity groups. Our results show that the larger models (such as ShieldGemma and LlamaGuard3) are comparatively more robust than the smaller BERT-style models due to being trained on large amounts of varied data sources. While none of the public CM models are completely robust, they are fairly robust across different ethnicity groups.