arXiv:2608.26013cs.CL2026-08

VISA通过自我进化循环生成更高质量的多模态指令数据。

VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following

论文配图:VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
图 1 · 摘自论文原文
  • 用智能体构建自进化循环,动态优化指令合成过程。
  • 在MM-IFEval上超越强基线,7个公开基准保持泛化能力。
  • 无需额外奖励模型,验证反馈直接用于强化学习。

多模态指令遵循模型需要准确、多样、可验证且具有挑战性的训练数据。现有合成流程通常采用单次生成-过滤范式,丢弃失败样本的反馈、验证结果和目标模型错误。本文提出VISA(视觉指令合成智能体),一种将多模态指令合成重构为自进化循环的智能体框架。每轮中,VISA分析图像以过滤不兼容约束并发现新可验证约束,从持久记忆中采样兼顾多样性和难度的约束集,生成候选指令,并通过可执行工具与结构化大语言模型裁判验证样本。失败样本触发诊断引导恢复,成功样本则被目标模型探测以估计难度。验证信号与目标模型失败特征被写回记忆,使后续轮次能自适应扩展约束空间、减少模板重复并聚焦未解决模型弱点。相同的验证合约进一步为强化学习提供奖励信号,无需独立训练的奖励模型。在MM-IFEval上的实验表明,VISA持续优于强基线,同时在七个公开基准上保持多模态通用能力。

原文摘要 · Abstract (English)

Multimodal instruction-following models require training data that is accurate, diverse, verifiable, and challenging. Existing synthesis pipelines typically follow a one-pass generate-and-filter paradigm, discarding feedback from failed samples, verifier outcomes, and target-model errors. We present VISA (Visual Instruction Synthesis Agent), an agentic framework that reformulates multimodal instruction synthesis as a self-evolving loop. At each round, VISA analyzes an image to filter incompatible constraints and discover new verifiable ones, samples diversity- and difficulty-aware constraint sets from persistent memory, generates candidate instructions, and verifies the resulting samples with executable tools and structured large language model judges. Failed samples trigger diagnostic-guided recovery, while accepted samples are probed against the target model to estimate difficulty. The resulting verifier signals and target-model failure profiles are written back to memory, allowing subsequent rounds to adaptively expand the constraint space, reduce template repetition, and focus on unresolved model weaknesses. The same verifier contracts further provide reward signals for reinforcement learning without a separately trained reward model. Experiments on MM-IFEval show that VISA consistently improves multimodal instruction following over strong baselines, while preserving general multimodal capability across seven public benchmarks.

多模态指令生成智能体自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。