根据用户偏好自动生成科研报告幻灯片,只需提供范例和模板。
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
- 用范例和模板隐式学习用户偏好,无需详细文字说明。
- 引入语音链机制,让幻灯片内容与演讲稿自然匹配。
- 适合需要快速生成个性化报告的科研人员或教学者。
自动幻灯片生成可大幅简化内容创作流程。然而,由于用户偏好各异,现有方法因需求不明确常导致结果不佳。本文提出新任务:基于用户指定偏好生成论文幻灯片。我们设计了受人类行为启发的智能体框架 SlideTailor,逐步生成可编辑的、符合用户偏好的幻灯片。系统仅需用户提供一篇论文-幻灯片示例对和视觉模板——这两种自然且易获取的输入,即可隐式编码丰富的内容与视觉风格偏好。尽管输入信息隐含且无标签,该框架仍能有效提炼并泛化偏好,指导个性化生成。此外,我们提出新颖的链式语音机制,使幻灯片内容与预设口头讲解同步,显著提升生成质量,并支持视频演示等下游应用。为推动此任务,我们构建了包含多样化用户偏好的基准数据集,并设计可解释的评估指标。大量实验验证了该框架的有效性。
原文摘要 · Abstract (English)
Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual user needs. We introduce a novel task that conditions paper-to-slides generation on user-specified preferences. We propose a human behavior-inspired agentic framework, SlideTailor, that progressively generates editable slides in a user-aligned manner. Instead of requiring users to write their preferences in detailed textual form, our system only asks for a paper-slides example pair and a visual template - natural and easy-to-provide artifacts that implicitly encode rich user preferences across content and visual style. Despite the implicit and unlabeled nature of these inputs, our framework effectively distills and generalizes the preferences to guide customized slide generation. We also introduce a novel chain-of-speech mechanism to align slide content with planned oral narration. Such a design significantly enhances the quality of generated slides and enables downstream applications like video presentations. To support this new task, we construct a benchmark dataset that captures diverse user preferences, with carefully designed interpretable metrics for robust evaluation. Extensive experiments demonstrate the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。