CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Abstract
Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, current models lack methods for eliciting such nuanced preferences and ground truth data to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios where information sharing violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants, given a communication role and AI-mediated condition. We formulate a prediction task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across eight open and proprietary models, personalization improves performance, with accuracy gains of up to 11.41% using six in-context examples. However, these gains arise from different mechanisms. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 leverage semantic context to infer user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability and present trade-offs between false positives and false negatives. Our findings highlight the potential and limitations of current LLMs' privacy modeling capability and position CIDER as a resource for advancing personalized privacy preference alignment.