arXiv:2412.12588cs.CL2024-12ACL被引 1

构建多视角信息检索与摘要框架,打破信息茧房

PerSphere: A Comprehensive Framework for Multi-Faceted Perspective Retrieval and Summarization

  • 设计双对立观点的文档检索与摘要流程
  • 现有模型在长文本和视角提取上表现不佳
  • 提出多智能体系统提升摘要全面性,适合研究者参考

随着在线平台和推荐算法的发展,用户越来越被困在信息回音室中,导致对各类议题的认知偏颇。为应对这一问题,我们提出了PerSphere,一个用于多维度视角检索与摘要的基准测试框架,旨在打破信息孤岛。每个查询包含两个对立主张,分别由一组不重叠的文档支撑,目标是准确总结这些文档,使摘要与对应主张及其背后视角一致。该任务采用端到端两步流程:全面文档检索与多视角摘要生成。此外,我们设计了一套评估指标来衡量检索与摘要内容的完整性。实验结果表明,当前模型在此复杂任务上表现有限。分析显示主要挑战在于长上下文处理和视角抽取,为此我们提出一种简单但高效的多智能体摘要系统,为提升PerSphere性能提供了可行方案。

原文摘要 · Abstract (English)

As online platforms and recommendation algorithms evolve, people are increasingly trapped in echo chambers, leading to biased understandings of various issues. To combat this issue, we have introduced PerSphere, a benchmark designed to facilitate multi-faceted perspective retrieval and summarization, thus breaking free from these information silos. For each query within PerSphere, there are two opposing claims, each supported by distinct, non-overlapping perspectives drawn from one or more documents. Our goal is to accurately summarize these documents, aligning the summaries with the respective claims and their underlying perspectives. This task is structured as a two-step end-to-end pipeline that includes comprehensive document retrieval and multi-faceted summarization. Furthermore, we propose a set of metrics to evaluate the comprehensiveness of the retrieval and summarization content. Experimental results on various counterparts for the pipeline show that recent models struggle with such a complex task. Analysis shows that the main challenge lies in long context and perspective extraction, and we propose a simple but effective multi-agent summarization system, offering a promising solution to enhance performance on PerSphere.

多视角信息茧房摘要生成评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。