通过模块化重构文档,实现多维度条件检索的精准训练。
Multi-Facet Blending for Faceted Query-by-Example Retrieval
- 将文档拆解为细粒度片段,利用大模型生成相关性对。
- 动态重组片段,构建面向特定维度的训练数据对。
- 无需预设标签,在教育试题检索中验证有效。
随着对细粒度用户意图需求的增长,基于特定维度的查询示例(Faceted QBE)检索受到关注。然而,以往方法主要依赖文献级比较,如引用关系,受限于缺乏细粒度相关性数据集,仅适用于引用型领域,难以捕捉维度约束的复杂性。本文提出多维度融合(FaBle)增强方法,通过分解与重组实现模块化合成:自动将文档拆分为维度单元,利用大模型内在区分能力生成相关/不相关对;动态重组单元,形成面向维度的相关性感知文档对。该方法无需预定义维度知识或标签。为进一步验证其在非引用领域的有效性,我们发布了用于教育试题检索的基准数据集。在1000篇文档上应用FaBle增强,显著提升了条件嵌入的训练效果。
原文摘要 · Abstract (English)
With the growing demand to fit fine-grained user intents, faceted query-by-example (QBE), which retrieves similar documents conditioned on specific facets, has gained recent attention. However, prior approaches mainly depend on document-level comparisons using basic indicators like citations due to the lack of facet-level relevance datasets; yet, this limits their use to citation-based domains and fails to capture the intricacies of facet constraints. In this paper, we propose a multi-facet blending (FaBle) augmentation method, which exploits modularity by decomposing and recomposing to explicitly synthesize facet-specific training sets. We automatically decompose documents into facet units and generate (ir)relevant pairs by leveraging LLMs' intrinsic distinguishing capabilities; then, dynamically recomposing the units leads to facet-wise relevance-informed document pairs. Our modularization eliminates the need for pre-defined facet knowledge or labels. Further, to prove the FaBle's efficacy in a new domain beyond citation-based scientific paper retrieval, we release a benchmark dataset for educational exam item QBE. FaBle augmentation on 1K documents remarkably assists training in obtaining facet conditional embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。