首个评估学术插图视觉逻辑一致性的基准,让AI生成的图更靠谱。
AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation
- 用VQA设计四层逻辑问题,逐级检验图与论文是否匹配。
- 实测显示模型间差距大,复杂推理和高密度生成能力差异明显。
- 逻辑与美学难兼顾,推理和生成能力提升可显著改善结果。
尽管图像生成技术快速发展,但当前顶尖模型能否生成可直接用于论文的学术插图仍缺乏研究。传统直接对比或依赖视觉语言模型(VLM)评估存在局限,尤其在长而复杂的文本与图表下可靠性不足。为此,我们提出AIBench,首个利用视觉问答(VQA)评估学术插图逻辑正确性、同时用VLM评估美学质量的基准。具体地,基于论文方法部分总结出的逻辑图,设计了四个层级的问题,从不同尺度检验生成插图与原文的一致性。该VQA方法在视觉-逻辑一致性评估上更准确、细致,且对评判VLM的能力依赖更低。借助高质量的AIBench,我们开展广泛实验,发现模型在此任务上的性能差距显著大于通用任务,反映出其在复杂推理与高密度生成上的能力差异。进一步实验表明,逻辑与美学难以同时优化,如同人工绘制;而测试时扩展推理与生成能力能显著提升表现。
原文摘要 · Abstract (English)
Although image generation has boosted various applications via its rapid evolution, whether the state-of-the-art models are able to produce ready-to-use academic illustrations for papers is still largely unexplored. Directly comparing or evaluating the illustration with VLM is native but requires oracle multi-modal understanding ability, which is unreliable for long and complex texts and illustrations. To address this, we propose AIBench, the first benchmark using VQA for evaluating logic correctness of the academic illustrations and VLMs for assessing aesthetics. In detail, we designed four levels of questions proposed from a logic diagram summarized from the method part of the paper, which query whether the generated illustration aligns with the paper on different scales. Our VQA-based approach raises more accurate and detailed evaluations on visual-logical consistency while relying less on the ability of the judger VLM. With our high-quality AIBench, we conduct extensive experiments and conclude that the performance gap between models on this task is significantly larger than general ones, reflecting their various complex reasoning and high-density generation ability. Further, the logic and aesthetics are hard to optimize simultaneously as in handcrafted illustrations. Additional experiments further state that test-time scaling on both abilities significantly boosts the performance on this task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。