arXiv:2603.21018cs.IRcs.AI2026-03

用可学习的领域语言统一结构化与非结构化数据检索,提升精准度。

DSL-R1: From SQL to DSL for Training Retrieval Agents across Structured and Unstructured Data with Reinforcement Learning

  • 将向量操作嵌入类SQL语法,融合逻辑推理与语义匹配
  • 在工业级邮件数据集上,命中率比基线高12.3%
  • 适合需要跨数据类型检索的智能搜索系统研发

复杂领域中的高效检索需弥合结构化元数据与非结构化内容之间的鸿沟。现有系统通常将两者割裂,依赖符号过滤或向量相似性,无法捕捉其相互作用。本文提出DSL-R1,一种通过新型领域特定语言(DSL)实现逻辑推理与语义匹配协同的统一框架。通过将向量原语嵌入类SQL操作符,该方法融合了符号计算的精确性与语义覆盖的广度。我们进一步引入强化学习机制,利用规则执行反馈与检索质量奖励联合优化DSL生成过程,平衡结构正确性与语义对齐性。在大规模工业级邮件基准上的评估表明,DSL-R1在Hit@1/3上实现了+12.3%的提升,持续优于解耦基线,确立了混合检索的新范式。

原文摘要 · Abstract (English)

Effective retrieval in complex domains requires bridging the gap between structured metadata and unstructured content. Existing systems typically isolate these capabilities, relying on either symbolic filtering or vector similarity, failing to capture their interplay. In this work, we propose DSL-R1, a unified framework that synergizes logical reasoning with semantic matching via a novel Domain-Specific Language (DSL). By embedding vector primitives within SQL-style operators, our approach leverages the complementary strengths of symbolic precision and semantic coverage. We further introduce a reinforcement learning mechanism where rule-based execution feedback and retrieval quality rewards jointly optimize the DSL generation, balancing structural correctness and semantic alignment. Evaluations on a large-scale industrial email benchmark demonstrate that DSL-R1 achieves a +12.3% improvement in Hit@1/3, consistently outperforming decoupled baselines and establishing a robust paradigm for hybrid retrieval.

检索增强领域语言强化学习混合检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。