用多个提示并行生成摘要,再融合提升质量。
Multi2: Multi-Agent Test-Time Scalable Framework for Multi-Document Processing
- 用多种提示并行生成候选摘要,再由聚合器整合
- 新指标有效缓解传统评估的位置偏差问题
- 适合需要高质量多文档摘要的场景
测试时扩展技术在提升大语言模型性能方面展现出潜力,通过推理阶段的战略性计算分配实现。尽管该方法在逻辑与数学推理任务中表现优异,但在自然语言生成(NLG)尤其是摘要任务中的应用仍待探索。多文档摘要(MDS)作为NLG的核心任务,需从多篇长文档中提取并整合关键信息,其挑战在于缺乏适用于所有需求的单一最优提示。为此,本文提出一种基于测试时扩展的MDS新框架:采用提示集成生成多个候选摘要,再通过聚合器生成优化结果。为有效评估,还引入两个基于LLM的新指标——一致性感知偏好(CAP)得分与LLM原子内容单元(LLM-ACU)得分,以缓解传统自动评估中的位置偏差。大量实验表明,该框架显著提升摘要质量,并揭示了应用于MDS的实际扩展边界。
原文摘要 · Abstract (English)
Recent advances in test-time scaling have shown promising results in improving Large Language Model (LLM) performance through strategic computation allocation during inference. While this approach has demonstrated strong improvements in logical and mathematical reasoning tasks, its application to natural language generation (NLG), particularly summarization, remains unexplored. Multi-Document Summarization (MDS), a fundamental task in NLG, presents unique challenges by requiring models to extract and synthesize essential information across multiple lengthy documents. Unlike reasoning tasks, MDS demands a more nuanced approach to prompt design and ensemble methods, as no single "best" prompt can satisfy diverse summarization requirements. We propose a novel framework leveraging test-time scaling for MDS. Our approach employs prompt ensemble techniques to generate multiple candidate summaries using various prompts, then combines them with an aggregator to produce a refined summary. To evaluate our method effectively, we also introduce two new LLM-based metrics: the Consistency-Aware Preference (CAP) score and LLM Atom-Content-Unit (LLM-ACU) score, which assess summary quality while addressing the positional bias inherent in traditional automatic evaluation. Our extensive experiments demonstrate that this framework significantly enhances summary quality while also revealing the practical scaling boundaries to MDS tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。