让视觉语言模型自动进化出更难、更智能的问题生成能力
Self-Evolving Visual Questioner

- 用模型自身生成并筛选问题,实现无监督自我提升
- 在相同预算下,生成问题的质量和难度显著优于静态数据训练
- 适合想提升模型主动探索能力的研究者或开发者
视觉语言模型(VLMs)通常作为被动回答者训练,其主动生成多样、非平凡、以视觉为中心且有依据的问题的能力尚未被充分探索。现有视觉提问系统的性能受限于高质量训练数据的可得性或标注成本。我们证明,VLM 可在无外部监督的情况下持续自我优化为视觉提问者。提出一种自演化框架,利用 VLM 自身作为问题提出者和过滤器,生成更难、更信息丰富、更具视觉中心性的提问,同时保持探索多样性以避免训练坍缩。这些问题被用于在提问者和回答者模式下训练 VLM。为评估提问能力,引入一种代理协议,从感知、推理和多样性维度评估问题质量。在多种骨干 VLM 上的实验表明,该方法显著提升了自主问题生成的质量,并大幅扩展了难度边界。在相同预算下,自监督比使用静态数据训练更有效。此外,自演化提问者在回答任务上仍保持竞争力甚至表现更优。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high-quality training data or the cost of curating them. We show that a VLM can continuously improve itself as a visual questioner without any external supervision. We propose a self-evolving framework that uses a VLM itself as both a proposer and a filter to produce harder, more informative, and visual-centric questions, while maintaining their exploration diversity to avoid training collapse. These questions are then used to train the VLM in both questioner and answerer modes. To evaluate the questioner, we introduce an agentic protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across various backbone VLMs show that our method substantially enhances the quality and substantially expands the difficulty boundary of autonomous question generation. Under the same budget, our self-supervision is more effective than training on the static source data. Moreover, the self-evolving questioner remains a competitive or even better answerer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。