arXiv:2606.00712cs.CV2026-06被引 1

基于Qwen的多模态推理系统在长时第一视角视频问答中夺冠

CASTLE2026 Team WDL Technical Report

论文配图:CASTLE2026 Team WDL Technical Report
图 1 · 摘自论文原文
  • 根据问题提示检索语音文本、图像和视频帧,按类型分发到专用提示模板
  • 通过多轮推理与置信度加权投票,最终得分0.58,显著优于基线(0.21)
  • 适合关注多模态融合与长时视频理解的研究者和开发者

CASTLE Challenge @ EgoVis 2026评估在600+小时多视角录制数据上的长时第一视角视频问答任务。每个四选一问题需结合视频、字幕、辅助图像、人物、日期、房间及时间上下文等证据。我们提出基于Qwen的证据感知多模态推理流程:解析问题线索,检索ASR片段,附加辅助图像,采样候选视频帧,并将问题路由至静态视觉、语音/文本、时间、混合四种类型,使用专用提示。通过多轮推理并以置信度加权投票聚合结果,最终转为Codabench官方格式。消融实验表明,使用LoRA使得分从0.21提升至0.50,增加采样帧数进一步提升至0.58。本系统在CASTLE Challenge @ EgoVis 2026中排名第一。

原文摘要 · Abstract (English)

The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-choice question requires evidence from videos, transcripts, auxiliary photos, people, days, rooms, and temporal context. We propose an evidence-aware multimodal reasoning pipeline based on Qwen. Our system parses question hints, retrieves ASR chunks, attaches auxiliary images, samples candidate video frames, and routes questions into static visual, speech/text, temporal, and mixed types with specialized prompts. Multiple inference passes are aggregated by confidence-weighted voting and converted into the official Codabench format. In ablation, LoRA improves the score from 0.21 to 0.50, and more sampled frames further raise it to 0.58. Our final system ranks first in the CASTLE Challenge @ EgoVis 2026.

多模态推理视频问答第一人称视频Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。