arXiv:2512.07538cs.CL2025-12中稿 · ACL被引 1

首个跨语言文档级语义差异识别数据集,助力多语言内容对齐评估

SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents

  • 构建跨语言文档级差异标注数据集,支持英德、英法、英意三组语言对
  • 模型在该任务上表现远低于单语句级基准,暴露现有模型能力短板
  • 适合研究多语言文本对齐、评测生成质量或改进跨语言理解的学者

识别文档间的语义差异对于文本生成评估和内容对齐至关重要,尤其在跨语言场景下。然而,这一独立任务尚未受到足够关注。为此,我们提出SwissGov-RSD,首个自然语境下的文档级跨语言语义差异识别数据集。该数据集包含224篇英-德、英-法、英-意多语言并行文档,由人工标注了细粒度的词级别差异。我们在该新基准上评估了多种开源与闭源大语言模型及编码器模型在不同微调设置下的表现。结果表明,当前自动方法在该任务上的性能显著低于其在单语句级及合成基准上的表现,揭示了大模型与编码器模型在跨语言语义差异识别方面存在明显差距。代码与数据集已公开。

原文摘要 · Abstract (English)

Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone task, it has received little attention. We address this by introducing SwissGov-RSD, the first naturalistic, document-level, cross-lingual dataset for semantic difference recognition. It encompasses a total of 224 multi-parallel documents in English--German, English--French, and English--Italian with token-level difference annotations by human annotators. We evaluate a variety of open-source and closed-source large language models as well as encoder models across different fine-tuning settings on this new benchmark. Our results show that current automatic approaches perform poorly compared to their performance on monolingual, sentence-level, and synthetic benchmarks, revealing a considerable gap for both LLMs and encoder models. We make our code and dataset publicly available.

跨语言语义差异评测基准文档对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。