arXiv:2605.27885cs.CV2026-05

通过师生对话机制,在推理时仅用少量标注数据实现视频问答的高效适应。

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering

  • 构建师生多轮对话,动态生成带反馈和视觉解释的上下文
  • 在EgoCross上超越零样本与直接上下文学习,获开源赛道第三
  • 无需微调,适合资源受限场景下的快速适配

现有方法如微调和上下文学习已被用于将视觉语言模型(VLM)适配至视频问答的特定领域。然而,在仅提供少量标注支持集的情况下,不经过微调而在推理阶段获取任务专属知识仍是挑战。本文提出一种仅依赖推理时上下文注入的方法。该方法首先构建一个反思对话(Reflective Dialogue, RD),由教师代理提出支持问题并给出正确性反馈,求解器代理作答并提供对正确与错误答案的视觉依据解释(即反思)。该对话历史作为推理时的上下文使用。在EgoCross基准上的实验表明,该方法优于基线零样本设置和标准的上下文学习方法,在CVPR 2026 EgoVis Workshop举办的首届跨域EgoCross挑战赛开源赛道中取得第三名,本文亦为此赛项的技术报告。

原文摘要 · Abstract (English)

Various approaches have been proposed to adapt Vision-Language Models (VLMs) to specialized domains for Video Question Answering, including fine-tuning and in-context learning. However, acquiring task-specific knowledge at the inference phase from only a small labeled support set without fine-tuning remains a challenge. In this paper, we propose a method that achieves adaptation solely through inference-time context injection. Our method first constructs a Reflective Dialogue (RD) -- a multi-turn conversation between two agents, in which Teacher poses each support question and delivers correctness feedback, and Solver answers and provides visual grounding explanations (or reflections) for both correct and incorrect answers. This dialogue history is then used as context at the inference phase. Experiments on the EgoCross benchmark demonstrate that our method outperforms both a baseline zero-shot setting and a standard in-context learning approach that passes support set examples directly, achieving 3rd place in the Open-source Track of the 1st Cross-Domain EgoCross Challenge at the CVPR 2026 EgoVis Workshop, for which this paper also serves as a technical report.

视频问答推理时学习对话机制少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。