发现手语翻译模型评估依赖签者,真实泛化能力远低于报告结果。
Rethinking Sign Language Translation: The Impact of Signer Dependence on Model Evaluation
- 用签者无关的交叉验证测试主流手语翻译模型性能
- 在PHOENIX14T上最高性能下降超80%,表明模型依赖签者特征
- 建议采用签者无关评估和句子不重叠数据划分以提升公平性
手语翻译虽借助深度学习取得进展,但评估仍高度依赖签者,训练、验证与测试集存在重叠签者。这引发疑问:模型是否真正具备泛化能力,还是仅依赖签者特有规律?我们在GFSLT-VLP、GASLT和SignCL三个主流无词素手语翻译模型上,对CSL-Daily和PHOENIX14T数据集进行签者交叉验证。在签者无关评估下,性能急剧下降:在PHOENIX14T上,GFSLT-VLP的BLEU-4从21.44降至3.59,ROUGE-L从42.49降至11.89;GASLT从15.74降至8.26;SignCL从22.74降至3.66。此外,在CSL-Daily中,许多目标句被多名签者表演,常见数据划分将相同句子同时置于训练与测试集,导致分数虚高,误判为泛化能力。研究揭示签者依赖评估会严重夸大手语翻译能力。建议:(1)采用签者无关协议以确保对未见签者的泛化;(2)重构数据集,提供显式的签者无关、句子不重叠划分;(3)同时报告签者依赖与签者无关结果,以及训练测试句重叠情况,以提升透明度与可比性。
原文摘要 · Abstract (English)
Sign Language Translation has advanced with deep learning, yet evaluations remain largely signer-dependent, with overlapping signers across train/dev/test. This raises concerns about whether models truly generalise or instead rely on signer-specific regularities. We conduct signer-fold cross-validation on GFSLT-VLP, GASLT, and SignCL, three leading, publicly available, gloss-free SLT models, on CSL-Daily and PHOENIX14T. Under signer-independent evaluation, performance drops sharply: on PHOENIX14T, GFSLT-VLP falls from BLEU-4 21.44 to 3.59 and ROUGE-L 42.49 to 11.89; GASLT from 15.74 to 8.26; and SignCL from 22.74 to 3.66. We also observe that in CSL-Daily many target sentences are performed by multiple signers, so common splits can place identical sentences in both training and test, inflating absolute scores by rewarding recall of recurring sentences rather than genuine generalisation. These findings indicate that signer-dependent evaluation can substantially overestimate SLT capability. We recommend: (1) adopting signer-independent protocols to ensure generalisation to unseen signers; (2) restructuring datasets to include explicit signer-independent, sentence-disjoint splits for consistent benchmarking; and (3) reporting both signer-dependent and signer-independent results together with train-test sentence overlap to improve transparency and comparability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。