arXiv:2503.09249cs.CLcs.AI2025-03NAACL被引 2

解决摘要生成中示例长度不一致的问题,提升效率与质量。

Considering Length Diversity in Retrieval-Augmented Summarization

  • 提出新算法DL-MMR,兼顾查询相关性与目标长度多样性。
  • 相比原MMR,内存节省78万倍,计算成本降低50万倍。
  • 适合需要高效生成多长度摘要的应用场景。

本研究针对检索增强型摘要生成中示例摘要长度多样性的影响进行深入分析,此前工作未涵盖此问题。提出一种长度感知的最大边际相关性(DL-MMR)算法,通过结合查询相关性与目标长度多样性,在检索增强摘要生成中实现更优的长度控制。不同于以往需对所有示例进行两两相关性比较的MMR方法,DL-MMR将示例目标长度纳入考量,避免示例间的相互比较,显著降低计算开销并节省内存。实验表明,相较于原始MMR算法,DL-MMR在保持信息量不变的前提下,内存使用减少781,513倍,计算成本降低500,092倍,验证了其有效性。

原文摘要 · Abstract (English)

This study investigates retrieval-augmented summarization by specifically examining the impact of exemplar summary lengths under length constraints, not covered by previous work. We propose a Diverse Length-aware Maximal Marginal Relevance (DL-MMR) algorithm to better control summary lengths. This algorithm combines the query relevance with diverse target lengths in retrieval-augmented summarization. Unlike previous methods that necessitate exhaustive exemplar exemplar relevance comparisons using MMR, DL-MMR considers the exemplar target length as well and avoids comparing exemplars to each other, thereby reducing computational cost and conserving memory during the construction of an exemplar pool. Experimental results showed the effectiveness of DL-MMR, which considers length diversity, compared to the original MMR algorithm. DL-MMR additionally showed the effectiveness in memory saving of 781,513 times and computational cost reduction of 500,092 times, while maintaining the same level of informativeness.

摘要生成检索增强长度控制高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。