Human vs Machine Translation Detection: A Cross-Model, Cross-Domain, and Low-Resource Analysis
Abstract
The growing reliance on web-crawled corpora for training large language models introduces a critical challenge: machine-translated (MT) content is often treated as human translation (HT), potentially degrading data quality and propagating synthetic artifacts, especially in low-resource settings. In this work, we study MT–HT detectability across high-resource (English, Spanish) and low-resource (Swahili, Afrikaans) languages using both decoder-only LLMs in zero-shot and few-shot settings and fine-tuned multilingual encoders trained at the sentence level and on token-chunked inputs. We find that decoder-only models provide limited discrimination overall, while encoder-based classifiers achieve high accuracy when evaluated on data generated by the same MT systems and translators used during training in a legislative domain, but fail to generalize to unseen translators and domains such as low-resource conversational text and medical content. To address this limitation, we evaluate models on longer inputs by chunking text into segments of up to 300 tokens, which substantially improves generalization, suggesting that sentence-level signals are insufficient for robust detection. Finally, we analyze the relationship between translation quality and detectability using COMET and BLEURT scores, showing that lower-quality MT is more easily identified, while high-quality MT remains difficult to distinguish from HT. Our results highlight the importance of input granularity and translation quality in building robust MT detection systems across domains and languages.