Failure-Aware Penetration Testing via Typed Failure Tokens: From Calibrated Prediction to Structured Recovery
Abstract
Large language models are increasingly used as penetration-testing copilots, but tool-interactive performance remains brittle: small command errors derail progress, recovery is inconsistent, and long-horizon context is easily lost. We present a failure-aware penetration-testing framework that makes error states explicit by requiring the model to emit typed failure tokens together with a structured action protocol. The model is trained to produce protocol-constrained outputs with planning, action, and completion tags, while conditioning on a compact trajectory memory and using training-time reasoning guidance plus unlikelihood loss to reduce spurious failure predictions. Evaluated on Cybench, the framework improves protocol compliance, step-wise decision quality, and subtask completion over standard fine-tuning baselines. Analysis further shows that failure-type prediction is not merely descriptive: representations become more separable under failure-aware training, and correcting the predicted failure type on confused cases leads to substantially better downstream recovery, indicating that explicit failure classification serves as a useful control signal for action selection.