通过多视角渐进聚合,提升长文档匹配的细节捕捉能力
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
- 基于文档子主题构建多个匹配视角,覆盖异质信息
- 采用渐进式时间聚合,有效融合不同视角特征
- 在新闻重复检测与法律案件检索中表现更优
长文档匹配旨在判断两篇文档的相关性,已广泛应用于各类场景。现有方法多采用层次化或长上下文模型,虽能实现粗粒度理解,但可能忽略细节。部分研究通过构建关于对齐子主题的相似句视图来聚焦细粒度匹配信号。然而,长文档通常包含多个子主题,匹配信号具有异质性。仅关注同源对齐子主题可能导致代表性不足,引发建模偏差。本文提出一种新框架,用于建模更具代表性的匹配信号:首先,通过文档对的子主题捕获多样化匹配信号;其次,基于子主题构建多个文档视图,以覆盖异质且有价值的细节;最后,针对现有空间聚合方法(如注意力机制)难以整合异质信息的问题,提出时间聚合策略,随训练进程逐步融合不同视图。实验结果表明,该学习框架在新闻重复检测与法律案件检索等任务上均表现出色。
原文摘要 · Abstract (English)
Long-form document matching aims to judge the relevance between two documents and has been applied to various scenarios. Most existing works utilize hierarchical or long context models to process documents, which achieve coarse understanding but may ignore details. Some researchers construct a document view with similar sentences about aligned document subtopics to focus on detailed matching signals. However, a long document generally contains multiple subtopics. The matching signals are heterogeneous from multiple topics. Considering only the homologous aligned subtopics may not be representative enough and may cause biased modeling. In this paper, we introduce a new framework to model representative matching signals. First, we propose to capture various matching signals through subtopics of document pairs. Next, We construct multiple document views based on subtopics to cover heterogeneous and valuable details. However, existing spatial aggregation methods like attention, which integrate all these views simultaneously, are hard to integrate heterogeneous information. Instead, we propose temporal aggregation, which effectively integrates different views gradually as the training progresses. Experimental results show that our learning framework is effective on several document-matching tasks, including news duplication and legal case retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。