ConformalGuard: False-Positive-Controlled Action Gating for LM Classifiers
Abstract
Language model classifiers are increasingly deployed for high-stakes decisions such as phishing screening, content moderation, and fraud triage, yet the gap between validation-set false-positive rates and deployment-time false-positive rates remains uncontrolled. A threshold tuned on a validation-set provides no finite-sample guarantee on the false-positive rate, and at realistic prevalence even a half-percentage-point miscalibration can double analyst workload. We present ConformalGuard, a model-agnostic conformal calibration layer that wraps any scoring function with finite-sample control over the rate at which legitimate inputs trigger each intervention tier. We evaluate classifiers at different scales: fine-tuned BERT and zero-shot Claude Sonnet 4 and GPT-4o, on public email corpora, comparing conformal calibration against raw thresholds, temperature scaling, and isotonic regression. Zero-shot models produce coarsely discretized scores (as few as 10 unique values) that render standard calibration methods ineffective, while conformal calibration consistently stays within the target false-positive rate. For Claude Sonnet 4, conformal calibration cuts the false-positive rate from 0.040 to 0.008. At 2% prevalence, this reduces unnecessary analyst reviews by 30%. The false-positive-rate guarantee holds under exchangeability; detection rates are supported empirically.