Spotlight让文档关键信息更抓人,用精选内容激发读者兴趣。
Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents
- 分两阶段训练模型:先微调,再用偏好优化对齐
- 相比传统摘要,显著提升可读性与用户参与度
- 适合需要快速吸引注意力的文本摘要场景
本文提出Spotlight,一种新型信息提取范式,通过突出文档中最引人注目的内容,生成简洁且具吸引力的叙事。与强调全面覆盖的传统摘要不同,Spotlight专注于激发读者深度参与。我们明确定义了Spotlight与相关概念的差异,并构建了专用数据集进行基准测试。为生成高质量Spotlight,我们采用两阶段方法:首先在基准数据上微调大语言模型,随后通过直接偏好优化(DPO)进行对齐。全面评估表明,该模型不仅能精准识别关键元素,还显著提升原文可读性与互动价值。
原文摘要 · Abstract (English)
In this paper, we introduce Spotlight, a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a document. Unlike traditional summaries, which prioritize comprehensive coverage, spotlights selectively emphasize intriguing content to foster deeper reader engagement with the source material. We formally differentiate spotlights from related constructs and support our analysis with a detailed benchmarking study using new datasets curated for this work. To generate high-quality spotlights, we propose a two-stage approach: fine-tuning a large language model on our benchmark data, followed by alignment via Direct Preference Optimization (DPO). Our comprehensive evaluation demonstrates that the resulting model not only identifies key elements with precision but also enhances readability and boosts the engagement value of the original document.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。