OpenStamp: A Watermark for Open-Source Language Models
Abstract
With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to LLMs and distinguishing it from human-written content. A prominent class of techniques embeds subtle but detectable signals in generated text by modifying token sampling probabilities. However, such methods are unsuitable for open-source models, where users have white-box access and can easily disable watermarking during inference. In this work, we introduce OpenStamp, a watermarking technique that implants detectable signals into the generated text by modifying just the final projection, or unembedding, layer. Through experiments across two models, we show that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior methods. The implanted watermarking signal is designed, and confirmed, to be more robust to paraphrasing attacks and is harder to scrub off through post-hoc fine-tuning. We share our code through an anonymized repository to enable developers to easily watermark their models, and additionally release watermarked versions of several popular open-source models produced using OpenStamp.