arXiv:2512.07738cs.CV2025-12

提升视频问答生成的精准与排序一致性

HLTCOE Evaluation Team at TREC 2025: VQA Track

  • 采用列表式学习框架,结合指针选择与加权损失优化答案排序
  • 在需时间推理和语义消歧的问题上准确率与稳定性显著提升
  • 适合关注多模态生成与排序融合的视觉语言研究者

HLTCOE评估团队参与了TREC VQA的答案生成(AG)任务。针对该任务,我们提出一种列表式学习框架,旨在提升答案生成的语义精度与排序一致性。给定一个视频-问题对,首先由基础多模态模型生成多个候选答案,随后使用一种基于新型带权重排名的掩码指针交叉熵损失训练的模型进行重排序。该目标函数融合了基于指针的候选选择、依赖排序的加权机制以及词汇受限下的掩码交叉熵,实现稳定且可解释的列表级优化。通过连接生成建模与判别式排序,该方法生成连贯且细粒度的答案列表。实验表明,在需要时间推理和语义消歧的问题上,该方法在准确率与排序稳定性方面均取得持续提升。

原文摘要 · Abstract (English)

The HLTCOE Evaluation team participated in TREC VQA's Answer Generation (AG) task, for which we developed a listwise learning framework that aims to improve semantic precision and ranking consistency in answer generation. Given a video-question pair, a base multimodal model first generates multiple candidate answers, which are then reranked using a model trained with a novel Masked Pointer Cross-Entropy Loss with Rank Weights. This objective integrates pointer-based candidate selection, rank-dependent weighting, and masked cross-entropy under vocabulary restriction, enabling stable and interpretable listwise optimization. By bridging generative modeling with discriminative ranking, our method produces coherent, fine-grained answer lists. Experiments reveal consistent gains in accuracy and ranking stability, especially for questions requiring temporal reasoning and semantic disambiguation.

视频问答多模态生成排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。