提升生成式检索在金融等复杂场景下的多步推理能力
Multi-Step Semantic Reasoning in Generative Retrieval
- 用分步提示+结构化指令增强模型推理链
- 在FinQA上检索准确率显著提升,一致性更强
- 适合需要深度语义推理的金融、法律等专业检索
生成式检索(GR)模型将文档库编码在参数中,直接为查询生成相关文档标识符。尽管该范式在检索任务中展现出潜力,但现有GR模型在涉及数值上下文的复杂查询(如财务报告中的语义推理)上表现不佳,受限于推理能力,导致检索精度不足,难以实用。为此,我们提出ReasonGR框架,旨在增强GR模型在数值语境下的多步语义推理能力。ReasonGR采用结合任务特定指令与分步推理引导的结构化提示策略,并引入聚焦推理的适配模块以优化推理相关参数的学习。在包含复杂文档的金融查询数据集FinQA上的实验表明,ReasonGR显著提升了检索准确率与一致性,显示出其在高推理需求检索场景中推动GR模型发展的潜力。
原文摘要 · Abstract (English)
Generative retrieval (GR) models encode a corpus within model parameters and generate relevant document identifiers directly for a given query. While this paradigm shows promise in retrieval tasks, existing GR models struggle with complex queries in numerical contexts, such as those involving semantic reasoning over financial reports, due to limited reasoning capabilities. This limitation leads to suboptimal retrieval accuracy and hinders practical applicability. We propose ReasonGR, a framework designed to enhance multi-step semantic reasoning in numerical contexts within GR. ReasonGR employs a structured prompting strategy combining task-specific instructions with stepwise reasoning guidance to better address complex retrieval queries. Additionally, it integrates a reasoning-focused adaptation module to improve the learning of reasoning-related parameters. Experiments on the FinQA dataset, which contains financial queries over complex documents, demonstrate that ReasonGR improves retrieval accuracy and consistency, indicating its potential for advancing GR models in reasoning-intensive retrieval scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。