arXiv:2607.09263cs.CV2026-07ACL

针对手语检索中视觉相似却语义不同的难题,提出基于视觉混淆的难负样本挖掘方法。

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

论文配图:Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval
图 1 · 摘自论文原文
  • 根据手语嵌入空间中的视觉混淆度构建难负样本,而非依赖语义相似性。
  • 在PHOENIX-2014T数据集上显著提升细粒度检索性能,同时保持粗粒度准确率。
  • 适合关注手语理解与多模态检索的开发者和研究者。

手语检索(SLRet)实现了对手语内容的高效访问,但在细粒度场景下仍表现脆弱,需区分视觉相似的手语动作。我们发现该问题并非源于模型容量不足,而是难负样本监督无效所致。具体而言,细粒度检索失败可归因于负样本分布不匹配:语义不同但视觉相似的手语动作极少被当作难负样本,而现有基于文本的挖掘策略无法捕捉此类视觉模糊性。为此,我们提出签注意识难负样本挖掘(SAN),在手语嵌入空间中基于视觉混淆度构造难负样本,而非依赖语言相似性。在PHOENIX-2014T数据集上的实验表明,SAN显著提升了细粒度检索性能,同时保持了粗粒度准确性,凸显了将负样本监督与视觉模糊性对齐的重要性。

原文摘要 · Abstract (English)

Sign Language Retrieval (SLRet) enables efficient access to sign language content but remains fragile in fine-grained scenarios where visually similar signs must be distinguished. We show that this limitation does not stem from model capacity, but from ineffective hard negative supervision. Specifically, we formulate fine-grained retrieval failures as a negative distribution mismatch: semantically distinct yet visually confusable signs are rarely treated as hard negatives, while existing text-based mining strategies fail to capture such visual ambiguity. To address this issue, we propose Sign-Aware Hard Negative Mining (SAN), which constructs hard negatives based on visual confusability in the sign embedding space rather than linguistic similarity. Experiments on PHOENIX-2014T demonstrate that SAN substantially improves fine-grained retrieval performance while preserving coarse-grained accuracy, highlighting the importance of aligning negative supervision with visual ambiguity in sign language retrieval.

手语检索负样本挖掘多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。