How Humans and LLMs Read Gender into “Gender-Neutral” Physical Descriptions
Abstract
Recent AI guidelines recommend replacing explicit gender labels with seemingly "objective" physical descriptions to mitigate representational harms. However, it remains unclear whether the linguistic descriptors of these traits carry their own latent associations of gender, mirroring the implicit gender inferences humans make from visual cues. We demonstrate that for both humans and Large Language Models (LLMs), seemingly neutral physical attributes (e.g., "a chiseled jawline") encode systematic, latent gender associations. To investigate this, we introduce a novel dataset of 316 unique physical attributes curated from diverse source domains and multiply annotated with human gender association ratings. Evaluating state-of-the-art LLMs reveals they only partially align with human judgments, exhibiting distinct biases such as compressed rating distributions and instability when evaluating non-binary identities. Notably, models that excel on explicit bias benchmarks show no advantage in capturing these implicit associations, and comparisons between base and instruction-tuned models reveal that post-training significantly alters this alignment. Finally, we train a specialized proxy model to predict human judgments of latent gender, enabling large-scale audits of the implicit gender encoded in descriptive language. Our results challenge the assumption that simply omitting explicit gender labels yields neutral text, exposing an underexplored dimension of potential implicit bias in language model safety and alignment.