让论文图稿通过多轮人机交互持续优化,效果显著优于现有方法。
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

- 用多智能体内部纠错-优化循环实现图稿多轮迭代改进。
- 相比基线系统,质量评分提升11.9至18.6分,遗忘率降低3.7至6.2点。
- 构建含3518条需求的基准数据集,支持可扩展的人机协作研究。
近期研究尝试从论文内容自动生成功能性科学图表(Lin等,2026;Zhu等,2026a)。然而单次生成难以完全满足作者的视觉与表达偏好:在一项预研用户研究中(N=14),所有参与者在查看初版后均要求进一步修改,86%认为优化后的图表更满意。尽管需求明确,多轮交互流程仍鲜有探索。为此,我们提出MTPaperBananaBench,一个包含292张图像和3,518条用户需求标注的多轮图表生成基准。为降低人工成本并实现可扩展评估,我们构建了一个用户模拟器,每轮识别未满足需求,并将其中k条转化为自然语言反馈。评估显示,现有基线多轮系统存在两类共性失败模式:(1)质量漂移,即图表质量随轮次递减;(2)遗忘,即先前实现的特征在后续轮次丢失。为应对这些问题,我们提出PaperBanana-Interact,一种通过内部批判-修正循环迭代优化图表的多智能体系统。该系统在多轮中保持甚至提升质量,相较基线质量得分提高11.9–18.6分,遗忘率降低3.7–6.2分。
原文摘要 · Abstract (English)
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。