arXiv:2508.19758cs.CLcs.IR2025-08EMNLP被引 6

让新闻检索更全面:通过细粒度句子分析避免信息重复

Uncovering the Bigger Picture: Comprehensive Event Understanding Via Diverse News Retrieval

  • 分两阶段检索:先找相关文本,再按句子聚类并重排以增强多样性
  • 在两个新构建的数据集上,多样性指标提升23%~38%,且不牺牲相关性
  • 适合需要多视角理解事件的研究者或新闻平台开发者

获取多元视角对理解现实事件至关重要,但多数新闻检索系统仅关注文本相关性,导致结果冗余、视角单一。我们提出NEWSCOPE,一种两阶段的多样化新闻检索框架,通过显式建模句子级别的语义差异来提升事件覆盖度。第一阶段使用稠密检索获取主题相关的内容,第二阶段采用句子级聚类与多样性感知重排序,挖掘互补信息。为评估检索多样性,我们引入三个可解释指标:平均成对距离、正向聚类覆盖率和信息密度比,并构建两个段落级基准数据集:LocalNews和DSGlobal。实验表明,NEWSCOPE持续优于强基线,在不降低相关性的前提下显著提升多样性。结果证明,细粒度、可解释的建模能有效缓解冗余,促进全面事件理解。数据与代码已公开于https://github.com/tangyixuan/NEWSCOPE。

原文摘要 · Abstract (English)

Access to diverse perspectives is essential for understanding real-world events, yet most news retrieval systems prioritize textual relevance, leading to redundant results and limited viewpoint exposure. We propose NEWSCOPE, a two-stage framework for diverse news retrieval that enhances event coverage by explicitly modeling semantic variation at the sentence level. The first stage retrieves topically relevant content using dense retrieval, while the second stage applies sentence-level clustering and diversity-aware re-ranking to surface complementary information. To evaluate retrieval diversity, we introduce three interpretable metrics, namely Average Pairwise Distance, Positive Cluster Coverage, and Information Density Ratio, and construct two paragraph-level benchmarks: LocalNews and DSGlobal. Experiments show that NEWSCOPE consistently outperforms strong baselines, achieving significantly higher diversity without compromising relevance. Our results demonstrate the effectiveness of fine-grained, interpretable modeling in mitigating redundancy and promoting comprehensive event understanding. The data and code are available at https://github.com/tangyixuan/NEWSCOPE.

新闻检索多样性事件理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。