arXiv:2607.27826cs.AIcs.CV2026-07被引 1

提出手语问答新任务,评估模型对手语视频的深层理解能力

Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

论文配图:Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
图 1 · 摘自论文原文
  • 设计手语问答任务,通过自然语言提问评估理解力
  • 构建两个基于PHOENIX14T和CSL-Daily的数据集,覆盖5类推理问题
  • 提出带条件时序下采样的基线模型,有效提升问答性能

近期手语理解(SLU)进展显著,但在连续手语识别和翻译任务上仍受限于预设目标,仅评估固定映射能力,难以检验模型对手语语义内容的真实理解。为此,本文首次提出新任务——手语问答(SLQA),要求模型回答关于手语视频的任意自然语言问题,以更灵活、全面地评估多维度推理能力。为支持该任务,我们基于PHOENIX14T和CSL-Daily构建了两个手语问答数据集(SignQA),利用精心设计的模板自动从现有词素和句子标注生成问答对,涵盖位置推理、结构推理、视觉搜索、词素识别和翻译理解五类互补问题。同时提出一种简单有效的基线模型,引入问题条件调制时序下采样模块和领域内知识迁移策略,实现从已有SLU任务的知识迁移,并增强对问题敏感的时序特征建模。大量实验表明,该基线在所有问题类别上均优于代表性视觉-语言模型,建立了可靠的基准。数据集已公开于:https://huggingface.co/datasets/hulala/SignQA-2026。

原文摘要 · Abstract (English)

Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-language sentences. As a result, they provide only a limited assessment of whether a model truly understands the semantic content of SL videos. To address this limitation, \textbf{we first propose a new task, Sign Language Question Answering (SLQA)}, which evaluates SL understanding by requiring models to answer arbitrary natural language questions about SL videos. Unlike previous SLU tasks, SLQA provides a more flexible and comprehensive evaluation framework that assesses multiple reasoning capabilities beyond recognition and translation. To facilitate this task, \textbf{we further construct two SignQA benchmarks} based on PHOENIX14T and CSL-Daily by automatically generating question-answer pairs from existing gloss and sentence annotations using carefully designed templates. The resulting datasets cover five complementary question categories, including position reasoning, structural reasoning, visual search, gloss recognition, and translation understanding. \textbf{Finally, we propose a simple yet effective baseline model} equipped with a Question-Conditioned Modulated Temporal Downsampling module and an in-domain knowledge transfer strategy, enabling effective knowledge transfer from existing SLU tasks while enhancing question-aware temporal feature modeling. Extensive experiments demonstrate that our baseline consistently outperforms representative vision-language models across all question categories, establishing a strong benchmark for future research on SLQA. Datasets are available at:{https://huggingface.co/datasets/hulala/SignQA-2026}.

手语理解问答系统多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。