Evaluating and Mitigating Misgendering in English-to-Hindi Machine Translation
Abstract
Machine translation between gender-neutral languages such as English and gender-marked Indian languages poses a fundamental challenge: systems must infer and realize grammatical gender in the target language, yet often default to masculine forms, rely on occupational stereotypes, or produce grammatically valid translations that obscure gender altogether. While gender bias in machine translation has been studied extensively for European languages, Indian language models remain under-evaluated, and existing benchmarks do not adequately capture the linguistic and cultural phenomena relevant to Indian settings. We introduce a 37,345-sentence benchmark for English-Hindi gender bias evaluation spanning twelve linguistically motivated categories, including stereotype inference, explicit gender preservation, Indian name disambiguation, counter-stereotype recognition, late-binding gender revelation, and Winograd-style coreference resolution. Paired with a grammar-aware Hindi gender classifier that combines rule-based morphological analysis with LLM fallback, the benchmark enables evaluation of four translation models: Helsinki Opus-MT, NLLB-200, Sarvam AI, and IndicTrans2. Across all four models, we find a consistent tendency to neutralize gender through Hindi’s ergative construction, with 63–76% of translations using syntax that obscures subject gender even when the English source contains explicit pronouns. Correspondingly, the explicit-gender control condition achieves only 17–31% classification accuracy, suggesting that explicit gender cues are often not preserved during translation. Resolved-prediction accuracy, our primary metric, measuring how often a model's gender prediction matches the intended referent after disambiguation, ranges from 47.8% for Helsinki Opus-MT to 59.1% for IndicTrans2, with all pairwise differences statistically significant (permutation test, p < 0.01). We further find that a stereotype-override prompt reduces masculine defaulting from 10.5% to 5.3% and improves resolved-prediction accuracy by 1.9 percentage points. These findings position Hindi as a qualitatively distinct setting for gender bias evaluation in machine translation and highlight the need for more context-aware mitigation strategies in Indian language NLP.