arXiv:2607.03006cs.CVcs.AI2026-07被引 3

用可审计框架评估科学海报生成模型是否遵守真实数据承诺。

PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

论文配图:PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark
图 1 · 摘自论文原文
  • 以占位符优先设计,分离布局与图像生成任务
  • 实测三篇论文合成图像率从34%降至0%
  • 适合评估生成模型在科研传播中的可信度

文本密集型图像模型已能设计海报级版式,但缺乏衡量其是否遵守科学传播规范的方法:文字可读、指定长宽比,以及最重要的是避免虚构科学图表。我们提出 POSTERHARNESS,一个可审计的框架,将海报生成重构为可量化的指令遵循任务,并建立初步基准与故障分类体系。POSTERHARNESS 采用占位符优先的契约机制,将模型职责拆分为视觉摘要设计(字体、阅读路径、色彩、背景)和不绘制数据图。每个图区必须是带标签的空白占位符;由确定性组合器在检测坐标处插入原始论文的图表。该设计使多项属性可测量:占位符数量与编号准确率、空白度、长宽比合规性、不生成合成图形、公开文本整洁度及源图溯源性——故障以显式拒绝形式记录,而非隐藏于看似合理的输出中。我们在12篇论文(6篇高能物理,6篇人工智能/机器学习相关)上实例化该框架,报告三项发现:(i) 反事实探测显示,占位符契约使三个论文的视觉语言模型识别出的合成图像数从34%降至0%;(ii) 故障分类识别出四大阻断性契约:占位符几何、占位符问答、模板批判与公共文本;(iii) 与 Paper2Poster 对比显示权衡:PosterHarness 生成更高分辨率作品,白底占比更低,且更受视觉语言模型偏好;确定性基线则稍多保留 PosterQuiz 风格信息,运行更快。我们将其报告为范式特征,而非优劣声明。所有成果物、提示词、清单文件与审计脚本均已开源,作为可复用评估组件。

原文摘要 · Abstract (English)

Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible labels, prescribed aspect ratios, and -- above all -- abstaining from fabricated scientific figures. We present POSTERHARNESS, an auditable harness reframing poster generation as measurable instruction-following tasks, with a pilot benchmark and failure taxonomy. POSTERHARNESS uses a placeholder-first contract to separate two jobs models otherwise conflate. The model performs visual-summary design: typography, reading path, color, and background -- but never draws data-bearing figures. Every figure region must be an empty labeled placeholder; a deterministic compositor inserts real source-paper figures at detected coordinates. This makes properties measurable: placeholder count and ID accuracy, blankness, aspect-ratio compliance, abstention from synthesized graphics, public-text hygiene, and source-figure provenance -- with failures logged as explicit rejections, not hidden in plausible-looking output. We instantiate the harness on 12 papers (6 HEP, 6 AI/ML-adjacent) and report three findings. (i) A counterfactual probe shows the placeholder contract drives VLM-counted synthesized figures from 34 to 0 across three papers. (ii) A failure taxonomy identifies blocking contracts: placeholder geometry, placeholder QA, template critic, and public text. (iii) Comparison with Paper2Poster shows a trade-off: PosterHarness yields higher-resolution artifacts, lower white-canvas fraction, and stronger VLM visual preference; the deterministic baseline retains slightly more PosterQuiz-style information and runs faster. We report this as regime characterization, not a superiority claim. All artifacts, prompts, manifests, and audit scripts are released as a reusable evaluation component.

海报生成可审计占位符VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。