arXiv:2510.05571cs.CL2025-10被引 13

用AI自动生成有故事性且美观的学术演讲,还能自我改进。

Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations

  • 构建可自我优化的智能演讲生成框架,融合叙事与美学设计。
  • 在650篇顶会论文上验证,能生成高质量含视觉与内容的演示材料。
  • 适合需要高效展示研究成果的研究者,尤其适合不擅长表达者。

学术论文推广已成为提升研究可见性的关键手段。然而现有自动化方法普遍存在叙事能力弱、美学质量不足、自我调整能力受限等问题,难以实现高效且吸引人的传播。核心挑战在于:无法准确评估,就无法有效改进。为此,我们提出EvoPresent——一个统一连贯叙事、美学感知设计与虚拟角色真实呈现的自进化代理框架。其核心是PresAesth,一种多任务强化学习美学模型,可提供可靠的美学评分、缺陷修正与对比反馈,支持在有限训练数据下实现迭代自我优化。为系统评估,我们构建EvoPresent基准,包含:基于650篇顶级会议论文的多模态资源(幻灯片、视频、脚本)的演示生成质量评估;以及2000对不同美学水平的幻灯片组成的美学感知评估,支持评分、缺陷修正与对比联合训练与测试。实验表明:(i) 高质量反馈对代理自我改进至关重要,初始能力不足以保证有效纠错;(ii) 自动化生成管道在视觉设计与内容构建间存在权衡;(iii) 多任务强化学习在美学感知任务中展现更强泛化能力。

原文摘要 · Abstract (English)

The promotion of academic papers has become an important means of enhancing research visibility. However, existing automated methods struggle limited storytelling, insufficient aesthetic quality, and constrained self-adjustment, making it difficult to achieve efficient and engaging dissemination. At the heart of those challenges is a simple principle: \emph{there is no way to improve it when you cannot evaluate it right}. To address this, we introduce \textbf{EvoPresent}, a self-improvement agent framework that unifies coherent narratives, aesthetic-aware designs, and realistic presentation delivery via virtual characters. Central to EvoPresent is \textbf{PresAesth}, a multi-task reinforcement learning (RL) aesthetic model that provides reliable aesthetic scoring, defect adjustment, and comparative feedback, enabling iterative self-improvement even under limited aesthetic training data. To systematically evaluate the methods, we introduce \textbf{EvoPresent Benchmark}, a comprehensive benchmark comprising: \textit{Presentation Generation Quality}, built on 650 top-tier AI conference papers with multimodal resources (slides, videos and scripts) to assess both content and design; and \textit{Aesthetic Awareness}, consisting of 2,000 slide pairs with varying aesthetic levels, supporting joint training and evaluation on scoring, defect adjustment, and comparison. Our findings highlight that (i) High-quality feedback is essential for agent self-improvement, while initial capability alone does not guarantee effective self-correction. (ii) Automated generation pipelines exhibit a trade-off between visual design and content construction. (iii) Multi-task RL training shows stronger generalization in aesthetic awareness tasks.

学术展示自进化美学生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。