为技术图像生成设计可量化的评估基准,精准检验科学图示的细节准确性。
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
- 构建基于评分标准的评测体系,分解科学图示正确性为6076项标准
- 主流模型在复杂技术图示上仅达0.801的评分准确率,显示显著差距
- 可反馈错误用于迭代优化,提升生成精度至0.865分
我们研究技术图像生成,即模型需从详细描述中合成信息密集、科学精确的图示,而非仅生成视觉可信的图像。为量化进展,提出TechImage-Bench,一个基于评分标准的基准,涵盖生物图谱、工程/专利图和通用技术图示。从真实教材和技术报告中收集654幅图,构建详细图像指令与多级评分标准,将正确性拆解为6,076个标准和44,131个二元校验。评分标准通过大模型从上下文文本和参考图中提取,并由基于大语言模型的自动判别器评估,采用系统性惩罚机制将子问题结果聚合为可解释的标准得分。在若干代表性文本到图像模型上进行评测发现,尽管开放域表现良好,最佳基线模型整体仅达0.801的评分准确率和0.576的准则得分,揭示细粒度科学保真度的巨大差距。最后,证明相同评分标准可提供有效监督:将失败校验反馈至编辑模型进行迭代优化,使强生成器的评分准确率从0.660提升至0.865,准则得分从0.382升至0.697。TechImage-Bench既提供了严谨的技术图像生成诊断工具,也为提升符合规范的科学图示生成提供了可扩展的信号。
原文摘要 · Abstract (English)
We study technical image generation, where a model must synthesize information-dense, scientifically precise illustrations from detailed descriptions rather than merely produce visually plausible pictures. To quantify the progress, we introduce TechImage-Bench, a rubric-based benchmark that targets biology schematics, engineering/patent drawings, and general technical illustrations. For 654 figures collected from real textbooks and technical reports, we construct detailed image instructions and a hierarchy of rubrics that decompose correctness into 6,076 criteria and 44,131 binary checks. Rubrics are derived from surrounding text and reference figures using large multimodal models, and are evaluated by an automated LMM-based judge with a principled penalty scheme that aggregates sub-question outcomes into interpretable criterion scores. We benchmark several representative text-to-image models on TechImage-Bench and find that, despite strong open-domain performance, the best base model reaches only 0.801 rubric accuracy and 0.576 criterion score overall, revealing substantial gaps in fine-grained scientific fidelity. Finally, we show that the same rubrics provide actionable supervision: feeding failed checks back into an editing model for iterative refinement boosts a strong generator from 0.660 to 0.865 in rubric accuracy and from 0.382 to 0.697 in criterion score. TechImage-Bench thus offers both a rigorous diagnostic for technical image generation and a scalable signal for improving specification-faithful scientific illustrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。