用数学推理数据训练智能体,让其更准找到当前决策瓶颈的解决策略。
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval

- 通过对比学习构建进展导向的嵌入空间,对齐具体情境与抽象规则。
- 在动态任务和静态技能检索中均超越现有模型,无需环境特训。
- 仅需公开推理数据即可训练,适合缺乏标注数据的智能体场景。
自我改进的智能体从过往轨迹中积累可复用的洞察,使洞察检索在将经验转化为行动指导方面变得日益重要。在每个决策步骤中,准确检索到能突破当前决策瓶颈的洞察,有助于智能体向目标推进,我们称此为“智能体洞察检索”。然而,现有方法主要依赖语义相似性建模,忽略了检索到的洞察是否真正解决当前困境。本文提出 InsightEmb,一种基于对比学习的嵌入框架,仅使用数学推理数据学习可迁移的进展导向检索几何结构。InsightEmb 同时学习将具体情境与抽象启发式规则对齐,并将具有相似进展结构的推理轨迹聚类。我们在动态智能体任务和静态技能检索基准上评估了 InsightEmb。无需任何环境特定训练,其在所有评估中均表现优于现有推理嵌入模型。结果表明,状态-洞察匹配的几何结构可跨领域迁移,使得仅通过公开可用的推理数据进行有效训练成为可能,无需昂贵的环境特定监督。
原文摘要 · Abstract (English)
Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated experience into actionable guidance. At each decision step, retrieving the right insight can help the agent progress toward its goal, a setting we refer to as agentic insight retrieval. However, existing retrieval methods primarily model semantic similarity, while overlooking whether a retrieved insight resolves the agent's current decision bottleneck. We propose InsightEmb, a contrastive embedding framework that learns transferable progress-oriented retrieval geometry using only mathematical reasoning data. InsightEmb jointly learns to align concrete situations with abstract heuristic rules and to cluster reasoning trajectories with similar progress structures. We evaluate InsightEmb on dynamic agent tasks and a static skill-retrieval benchmark. Without any environment-specific training, InsightEmb improves over all these evaluations, surpassing the performance of existing reasoning embedding models. These results suggest that the geometry of state-insight matching can transfer across domains, enabling effective training from publicly available reasoning data without expensive environment-specific supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。