构建房地产问答基准,支持多步骤推理与工具调用。
ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

- 设计包含29,270个实例的基准数据集,支持中间步骤验证。
- 提出HIRE-Agent框架,实现理解-规划-执行的分层协作。
- 验证分层架构对复杂现实任务的重要性,适合智能客服研究者。
构建可导航碎片化、多源信息的智能体仍具挑战性,主要因缺乏结合数据库查询与外部API的混合工作流基准。为此,我们提出ReCoQA,一个包含29,270个房地产实例的大规模基准,提供可机器验证的中间步骤监督,包括结构化意图标签、SQL查询和API调用。同时,我们提出HIRE-Agent,一种分层框架,采用理解-规划-执行架构作为强基线。通过前端解析器、规划调度器和执行专家的协同,有效整合异构证据。大量实验表明,HIRE-Agent构成有力基线,证实了复杂现实推理任务中分层协作的必要性。
原文摘要 · Abstract (English)
Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting hybrid workflows combining database querying with external APIs. To bridge this gap, we introduce ReCoQA, a large-scale benchmark of 29,270 real-estate instances featuring machine-verifiable supervision for intermediate steps, including structured intent labels, SQL queries, and API calls. Complementarily, we propose HIRE-Agent, a hierarchical framework instantiating an understand-plan-execute architecture as a strong baseline. By orchestrating a Front-end parser, a planning Supervisor, and execution Specialists, HIRE-Agent effectively integrates heterogeneous evidence. Extensive experiments demonstrate that HIRE-Agent constitutes a strong baseline and substantiates the necessity of hierarchical collaboration for complex, real-world reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。