首个支持印地语族多任务评估的质检与自动校对基准,统一九组语言对数据。
IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages
- 整合2020-2024年WMT任务数据,构建含四种标签的统一数据集
- 发现质量信号冲突的片段评分显著更低,且此现象跨模型跨语言一致
- 适用于机器翻译质检、自动校对及多语言模型评估的研究者
印地语族质量估计(QE)与自动后编辑(APE)数据分散于多个发布版本中,缺乏统一资源支持跨任务、跨语言对的训练与评估。本文将WMT 2020–2024共享任务谱系与扩展的英-马拉雅拉姆语数据合并为IndicQE:共126,754个实例,覆盖九个方向语言对,每条数据段对齐多达四种标签——直接评分、人工后编辑、词级正确/错误标记及错误解释,并构建按四个难度轴分层的测试集。在该数据集上,我们评估六种提示式大模型和三种COMET度量在段级质量估计的表现,以及三种系统在自动后编辑上的表现。其中两个难度轴基于直接评分定义并选取其压缩子集,每个轴均与同语言对、同评分分布的对照组比较。仅有一个轴通过检验:整体与词级质量信号冲突的片段,其评分显著低于同等得分的其他片段,且该现象在全部九个模型与七组语言对中一致。标注者分歧在引入对照组后不再产生影响。少样本提示使所有模型相关性与输出格式合规性下降不超过3.4B。同一语言内准确率无法保证跨语言可比性:三类训练度量中,内部相关性最佳者在跨语言合并时性能损失最严重。基准数据集(https://huggingface.co/datasets/surrey-nlp/IndicQE-APE)与代码(https://github.com/surrey-nlp/IndicQE-APE)已开源。
原文摘要 · Abstract (English)
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes. On it we benchmark six prompted LLMs and three COMET metrics on segment-level QE, and three systems on APE. Two of the axes are defined partly on the direct assessment and select a compressed slice of it, so each axis is compared against a control drawn from the same language pair with the same score distribution. Only one survives that control: segments whose holistic and token-level quality signals conflict are ranked worse than equally-scored segments of the same language, for all nine systems and all seven pairs that carry the axis. Annotator disagreement, which looks second-hardest without the control, has no effect with it. Few-shot prompting costs every model $\leq$ $3.4$B both correlation and output-format compliance. Within-language accuracy does not make scores comparable across pairs: of the three trained metrics, the one with the best within-language correlation loses most when the pairs are pooled. The benchmark (https://huggingface.co/datasets/surrey-nlp/IndicQE-APE) and code (https://github.com/surrey-nlp/IndicQE-APE) are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。