arXiv:2607.18225econ.EMcs.LG2026-07

用向量搜索匹配因果推断中的近邻,提升策略学习效果

Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference

  • 通过向量搜索找相似样本,结合生成模型预测结果
  • 两步法可证明后悔值有界,且依赖最近邻估计器性能
  • 适合需要可解释策略的因果推断场景

我们提出一种基于检索增强生成(RAG)的一步和两步策略学习方法。在潜在结果框架下,将RAG动作选择建模为近邻匹配问题。两步法中,向量搜索在嵌入空间中检索特定动作的邻近证据,生成器估计条件期望结果或其差异,再通过插件规则选择动作。该方法将动作特定的向量搜索与因果推断中的最近邻匹配相联系。我们将两步法的后悔值分解为候选生成后悔和候选内选择后悔,并利用最近邻估计器和Transformer的预测误差保证来界定后者。一步法因其中间计算不可观测,直接作为策略进行评估。

原文摘要 · Abstract (English)

We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedding space, the generator estimates conditional expected outcomes or their contrasts, and a plug-in rule selects an action. This formulation connects action-specific vector search with nearest-neighbor matching in causal inference. We decompose the regret of the two-step method into candidate-generation regret and within-candidate choice regret, and we bound the latter using prediction-error guarantees for nearest-neighbor estimators and transformers. We evaluate the one-step method directly as a policy because its intermediate computation is unobserved.

因果推断RAG策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。