评测意大利语无性别重写模型,发现开源模型表现更优且可小型化。
Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- 构建双维度评估框架,同时衡量无性别化与语义保真度。
- 微调后模型性能媲美大模型,参数量仅为几分之一。
- 揭示训练数据优化中无性别与语义保留的权衡问题。
无性别重写(GNR)旨在消除文本中不必要的性别标识,同时保持原意,这对意大利语等具有语法性别的语言尤为挑战。本文首次系统评估了当前最先进的大语言模型(LLMs)在意大利语GNR任务上的表现,提出一个二维评估框架,兼顾中立性与语义保真度。我们对比了多种大模型的少样本提示策略,对选定模型进行微调,并应用针对性清洗提升任务相关性。结果表明,开源权重模型优于唯一现有专用意大利语GNR模型;而我们的微调模型在性能上可匹配甚至超越最佳开源模型,但仅需其极小参数量。最后,我们讨论了训练数据优化中追求无性别化与语义保留之间的权衡。
原文摘要 · Abstract (English)
Gender-neutral rewriting (GNR) aims to reformulate text to eliminate unnecessary gender specifications while preserving meaning, a particularly challenging task in grammatical-gender languages like Italian. In this work, we conduct the first systematic evaluation of state-of-the-art large language models (LLMs) for Italian GNR, introducing a two-dimensional framework that measures both neutrality and semantic fidelity to the input. We compare few-shot prompting across multiple LLMs, fine-tune selected models, and apply targeted cleaning to boost task relevance. Our findings show that open-weight LLMs outperform the only existing model dedicated to GNR in Italian, whereas our fine-tuned models match or exceed the best open-weight LLM's performance at a fraction of its size. Finally, we discuss the trade-off between optimizing the training data for neutrality and meaning preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。