Researchers release GAND, a benchmark to study gender bias in machine translation through gender-ambiguous natural data
The GAND resource provides English source sentences designed to analyze how machine translation systems handle gender in the absence of clear cues, enabling interpretability and attribution analyses.
1 source · cross-referenced
- Researchers introduced GAND, a benchmarking resource for machine translation focused on gender-ambiguous natural data.
- GAND consists of English source sentences to study how contextual cues influence gender translation in grammatical gender languages.
- The resource enables interpretability and feature attribution analyses to identify source words influencing gendered translations.
- The authors demonstrate GAND by translating a subset into two grammatical gender languages and adding contrastive translations.
- The work is accepted for presentation at the EAMT 2026 technical track.
Researchers from multiple institutions have introduced GAND, a benchmarking resource designed to study gender bias in machine translation through gender-ambiguous natural data. The resource consists of English source sentences specifically crafted to analyze how machine translation systems handle gender when clear cues are absent. This focus addresses a gap in current evaluation methods, which often overlook scenarios where gender is ambiguous but consequential for translation quality and fairness.
The authors leverage GAND to conduct an interpretability analysis by translating a subset of the data into two grammatical gender languages. They extend these translations with manually crafted contrastive examples, enabling a closer examination of how contextual cues influence gendered translations. This approach allows researchers to isolate the source words and phrases that drive gendered choices in the target language, providing actionable insights into model behavior.
The resource is intended to support feature attribution analysis, a technique that identifies which parts of the input most influence the model's output. By applying this method to gender-ambiguous translations, the authors aim to reveal patterns of bias and default behavior in machine translation systems. This work is accepted for presentation at the EAMT 2026 technical track, indicating its relevance to the machine translation research community.
The motivation behind GAND stems from the observation that machine translation systems frequently produce gender-biased translations, often defaulting to stereotypes or oversimplified assumptions about gender. In contexts where self-expression and accuracy are paramount, such mistranslations can lead to harm for users who rely on these systems. GAND provides a structured way to diagnose these issues and develop more equitable translation technologies.
- Jul 28, 2026 · arXiv cs.CL
Official conference reviewer guidelines outperform LLM-generated reviewer-imitating guidelines in automated peer review study
Trust79 - Jul 28, 2026 · arXiv cs.CL
Researchers release MioFFAn, an open-source framework for annotating and formalizing scientific formulas with LLM-assisted workflows
Trust79 - Jul 28, 2026 · arXiv cs.AI
Researchers propose SeT-Diff, a diffusion-based foundational model for HPC telemetry and time-series
Trust79