构建金融中文段落检索的多层级查询生成数据集
FinCPRG: A Bidirectional Generation Pipeline for Hierarchical Queries and Rich Relevance in Financial Chinese Passage Retrieval
- 双向生成:从文档内句子到跨文档标题分层构造查询
- 覆盖1300份报告,生成含三层结构的查询与丰富相关性标签
- 适合金融信息检索与大模型训练,支持跨文档理解
近年来,大语言模型在构建段落检索数据集方面展现出巨大潜力。然而,现有方法在表达跨文档查询需求和控制标注质量方面仍存在局限。为此,本文提出一种双向生成管道,旨在为文档内与跨文档场景生成三级层次化查询,并在直接映射标注基础上挖掘额外相关性标签。该管道引入两种查询生成方法:自底向上从单文档文本生成句级与段落级结构化查询;自顶向下结合行业、主题、时间三个关键金融要素,将报告标题聚类后生成主题级查询。在相关性标注方面,不仅依赖生成关系的直接映射,还采用间接正例挖掘方法扩充相关查询-段落对。基于此管道,我们从近1300份中文金融研究报告中构建了金融段落检索生成数据集(FinCPRG),包含层次化查询与丰富的相关性标签。通过挖掘相关性标签评估、基准测试与训练实验,验证了FinCPRG在训练与评测中的有效性。
原文摘要 · Abstract (English)
In recent years, large language models (LLMs) have demonstrated significant potential in constructing passage retrieval datasets. However, existing methods still face limitations in expressing cross-doc query needs and controlling annotation quality. To address these issues, this paper proposes a bidirectional generation pipeline, which aims to generate 3-level hierarchical queries for both intra-doc and cross-doc scenarios and mine additional relevance labels on top of direct mapping annotation. The pipeline introduces two query generation methods: bottom-up from single-doc text and top-down from multi-doc titles. The bottom-up method uses LLMs to disassemble and generate structured queries at both sentence-level and passage-level simultaneously from intra-doc passages. The top-down approach incorporates three key financial elements--industry, topic, and time--to divide report titles into clusters and prompts LLMs to generate topic-level queries from each cluster. For relevance annotation, our pipeline not only relies on direct mapping annotation from the generation relationship but also implements an indirect positives mining method to enrich the relevant query-passage pairs. Using this pipeline, we constructed a Financial Passage Retrieval Generated dataset (FinCPRG) from almost 1.3k Chinese financial research reports, which includes hierarchical queries and rich relevance labels. Through evaluations of mined relevance labels, benchmarking and training experiments, we assessed the quality of FinCPRG and validated its effectiveness as a passage retrieval dataset for both training and benchmarking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。