arXiv:2512.15006cs.CVcs.AI2025-12中稿 · ed

评估视频问答生成问题质量,挖掘专家未知知识。

Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation

  • 用问答检索模拟专家对话,评估问题质量。
  • 在27,666个真实问答对上验证,上下文越丰富得分越高。
  • 适合研究知识挖掘与智能访谈系统的人看。

熟练的访谈者能从专家处提取宝贵信息,这引出核心问题:哪些问题更有效?为回答此问题,需对问题生成模型进行量化评估。视频问答生成(VQG)是视频问答领域的一项任务,即针对给定答案生成问题。现有评估多关注回答能力,而非生成问题的质量。本文聚焦于问题质量在挖掘人类专家未知知识方面的表现。为持续优化VQG模型,我们提出一种协议:通过问答检索模拟与专家的交流过程。为此构建新数据集EgoExoAsk,包含27,666个由Ego-Exo4D专家注释生成的问答对。使用训练集训练检索器,基准测试基于验证集上的Ego-Exo4D视频片段构建。实验结果表明,该度量与问题生成设置合理一致:能利用更丰富上下文的模型得分更高,验证了协议有效性。数据集已开源:https://github.com/omron-sinicx/VQG4ExpertKnowledge。

原文摘要 · Abstract (English)

Skilled human interviewers can extract valuable information from experts. This raises a fundamental question: what makes some questions more effective than others? To address this, a quantitative evaluation of question-generation models is essential. Video question generation (VQG) is a topic for video question answering (VideoQA), where questions are generated for given answers. Their evaluation typically focuses on the ability to answer questions, rather than the quality of generated questions. In contrast, we focus on the question quality in eliciting unseen knowledge from human experts. For a continuous improvement of VQG models, we propose a protocol that evaluates the ability by simulating question-answering communication with experts using a question-to-answer retrieval. We obtain the retriever by constructing a novel dataset, EgoExoAsk, which comprises 27,666 QA pairs generated from Ego-Exo4D's expert commentary annotation. The EgoExoAsk training set is used to obtain the retriever, and the benchmark is constructed on the validation set with Ego-Exo4D video segments. Experimental results demonstrate our metric reasonably aligns with question generation settings: models accessing richer context are evaluated better, supporting that our protocol works as intended. The EgoExoAsk dataset is available in https://github.com/omron-sinicx/VQG4ExpertKnowledge .

视频问答知识挖掘评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。