arXiv:2604.05514cs.AI2026-04ACL被引 3

统一生成各类图表代码,用视觉提问提升生成质量

OmniDiagram: Advancing Unified Diagram Code Generation via Visual Interrogation Reward

论文配图:OmniDiagram: Advancing Unified Diagram Code Generation via Visual Interrogation Reward
图 1 · 摘自论文原文
  • 通过视觉提问机制动态评估图表生成效果
  • 在196,000条数据上训练,性能达新SOTA
  • 无需人工标注真值代码,适合多类型图表生成

可编程图表生成正快速发展,对结构化可视化至关重要。但现有研究多局限于少数任务形式和语言支持,难以适配多样图表类型。本文提出OmniDiagram统一框架,融合多种图表代码语言与任务定义。为解决强化学习中代码逻辑与视觉一致性难题,提出新型视觉反馈策略Visual Interrogation Verifies All(Viva)。不同于脆弱的语法规则或像素级匹配,Viva通过生成针对性视觉问题来审查渲染图表的视觉保真度,并提供细粒度反馈以优化生成。该机制实现自进化训练,无需人工标注真值代码。此外,构建首个大规模图表代码生成数据集M3²Diagram,包含超过196,000个高质量样本。实验表明,结合监督微调与基于Viva的强化学习,OmniDiagram在多个图表代码生成基准上达到新SOTA。

原文摘要 · Abstract (English)

The paradigm of programmable diagram generation is evolving rapidly, playing a crucial role in structured visualization. However, most existing studies are confined to a narrow range of task formulations and language support, constraining their applicability to diverse diagram types. In this work, we propose OmniDiagram, a unified framework that incorporates diverse diagram code languages and task definitions. To address the challenge of aligning code logic with visual fidelity in Reinforcement Learning (RL), we introduce a novel visual feedback strategy named Visual Interrogation Verifies All (\textsc{Viva}). Unlike brittle syntax-based rules or pixel-level matching, \textsc{Viva} rewards the visual structure of rendered diagrams through a generative approach. Specifically, \textsc{Viva} actively generates targeted visual inquiries to scrutinize diagram visual fidelity and provides fine-grained feedback for optimization. This mechanism facilitates a self-evolving training process, effectively obviating the need for manually annotated ground truth code. Furthermore, we construct M3$^2$Diagram, the first large-scale diagram code generation dataset, containing over 196k high-quality instances. Experimental results confirm that the combination of SFT and our \textsc{Viva}-based RL allows OmniDiagram to establish a new state-of-the-art (SOTA) across diagram code generation benchmarks.

图表生成强化学习自进化数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。