用大模型生成金融问答数据,构建更贴近真实银行场景的检索测试集。
Query Generation Pipeline with Enhanced Answerability Assessment for Financial Information Retrieval
- 通过大模型自动生成单文档与多文档查询,实现低成本建模。
- 新评估方法比现有方法更贴近人工判断,准确率提升12.3%。
- 适合研究金融领域信息检索、大模型应用的开发者和研究人员。
随着大语言模型在金融领域的应用日益受关注,准确的信息检索(IR)对可靠AI服务至关重要。然而,现有基准无法反映真实银行场景中复杂且特定领域的信息需求。构建领域专属的IR基准成本高昂,且受限于客户数据的法律合规问题。为此,我们提出一种基于大模型的系统化方法,用于生成领域特定的IR基准。作为该方法的具体实现,我们的管道结合了单文档与多文档查询生成,并引入增强型推理辅助答案可回答性评估方法,在与人类判断的一致性上优于先前方法。基于此方法,我们构建了KoBankIR数据集,包含815个查询,源自204份官方银行文档。实验表明,现有检索模型在KoBankIR中的复杂多文档查询上表现不佳,验证了该系统化方法在构建领域基准方面的价值,并凸显了金融领域需改进检索技术的紧迫性。
原文摘要 · Abstract (English)
As financial applications of large language models (LLMs) gain attention, accurate Information Retrieval (IR) remains crucial for reliable AI services. However, existing benchmarks fail to capture the complex and domain-specific information needs of real-world banking scenarios. Building domain-specific IR benchmarks is costly and constrained by legal restrictions on using real customer data. To address these challenges, we propose a systematic methodology for constructing domain-specific IR benchmarks through LLM-based query generation. As a concrete implementation of this methodology, our pipeline combines single and multi-document query generation with an enhanced and reasoning-augmented answerability assessment method, achieving stronger alignment with human judgments than prior approaches. Using this methodology, we construct KoBankIR, comprising 815 queries derived from 204 official banking documents. Our experiments show that existing retrieval models struggle with the complex multi-document queries in KoBankIR, demonstrating the value of our systematic approach for domain-specific benchmark construction and underscoring the need for improved retrieval techniques in financial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。