针对意大利语性别刻板印象问题,提出三任务评估框架提升语言公平性。
GFG -- Gender-Fair Generation: A CALAMITA Challenge
- 构建三任务评估体系:检测、改写与跨语言生成性别公平表达
- 使用BERTScore、中性分类器等指标,量化评估性别公平性表现
- 专为意大利语设计数据集,支持非二元性别新词的公平表达测试
性别公平语言旨在通过包容性术语和表达促进性别平等,避免强化性别刻板印象。在意大利语等高度性别标记的语言中实现这一目标尤为困难。为此,本文提出性别公平生成挑战,旨在推动书面沟通中的性别公平语言发展。该挑战涵盖三个任务:(1)检测意大利语句子中的性别化表达;(2)将性别化表达改写为性别公平替代形式;(3)在英译意自动翻译中生成性别公平语言。挑战基于三个标注数据集:来自布雷西亚大学行政文档的GFL-it语料库;基于Europarl子集构建的双语测试集GeNTE,用于中性重写与翻译;以及面向非二元性别新词的双语测试集Neo-GATE,用于公平表达与翻译评估。每项任务采用特定评估指标:任务1使用BERTScore计算每个条目的平均F1分数;任务2采用性别中性分类器准确率;任务3采用加权覆盖准确率。
原文摘要 · Abstract (English)
Gender-fair language aims at promoting gender equality by using terms and expressions that include all identities and avoid reinforcing gender stereotypes. Implementing gender-fair strategies is particularly challenging in heavily gender-marked languages, such as Italian. To address this, the Gender-Fair Generation challenge intends to help shift toward gender-fair language in written communication. The challenge, designed to assess and monitor the recognition and generation of gender-fair language in both mono- and cross-lingual scenarios, includes three tasks: (1) the detection of gendered expressions in Italian sentences, (2) the reformulation of gendered expressions into gender-fair alternatives, and (3) the generation of gender-fair language in automatic translation from English to Italian. The challenge relies on three different annotated datasets: the GFL-it corpus, which contains Italian texts extracted from administrative documents provided by the University of Brescia; GeNTE, a bilingual test set for gender-neutral rewriting and translation built upon a subset of the Europarl dataset; and Neo-GATE, a bilingual test set designed to assess the use of non-binary neomorphemes in Italian for both fair formulation and translation tasks. Finally, each task is evaluated with specific metrics: average of F1-score obtained by means of BERTScore computed on each entry of the datasets for task 1, an accuracy measured with a gender-neutral classifier, and a coverage-weighted accuracy for tasks 2 and 3.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。