arXiv:2511.14670cs.AI2025-11

通过挖掘高价值动作构建领域技能图,提升上下文学习的决策效果。

SkillGen: Learning Domain Skills for In-Context Sequential Decision Making

  • 基于轨迹构建领域级动作图,自动识别关键决策步骤。
  • 在多个任务上平均提升进度率5.9%至16.5%,效果稳定。
  • 适合需要精准序列决策的智能体任务,如游戏导航与科学推理。

大语言模型(LLM)通过上下文学习(ICL)应用于序列决策,但其性能高度依赖提示质量。有效提示需满足三原则:聚焦决策关键信息、提供步骤级粒度、通过标签效率降低对专家标注的依赖。现有方法常无法同时满足三者。为此,我们提出SkillGen,一种面向结构化序列推理的基于技能的ICL框架。该框架从采样轨迹构建以动作为中心的领域级图,通过时序差分信用分配识别高价值动作,并检索步骤级技能生成细粒度、上下文感知的提示。我们进一步给出理论分析,表明聚焦高价值片段有助于任务可辨识性,指导更优的ICL提示设计。在ALFWorld、BabyAI和ScienceWorld上的实验显示,使用开源与专有LLM均取得一致提升,平均进度率提高5.9%–16.5%。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly applied to sequential decision-making through in-context learning (ICL), yet their effectiveness is highly sensitive to prompt quality. Effective prompts should meet three principles: focus on decision-critical information, provide step-level granularity, and minimize reliance on expert annotations through label efficiency. However, existing ICL methods often fail to satisfy all three criteria simultaneously. Motivated by these challenges, we introduce SkillGen, a skill-based ICL framework for structured sequential reasoning. It constructs an action-centric, domain-level graph from sampled trajectories, identifies high-utility actions via temporal-difference credit assignment, and retrieves step-wise skills to generate fine-grained, context-aware prompts. We further present a theoretical analysis showing that focusing on high-utility segments supports task identifiability and informs more effective ICL prompt design. Experiments on ALFWorld, BabyAI, and ScienceWorld, using both open-source and proprietary LLMs, show that SkillGen achieves consistent gains, improving progress rate by 5.9%-16.5% on average across models.

序列决策上下文学习技能挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。