arXiv:2502.17538cs.CL2025-02

用因果方法在自然语言动作空间中高效学习决策策略。

Policy Learning with a Natural Language Action Space: A Causal Approach

  • 单模型Q-learning优化语言嵌入,实现数据高效政策学习。
  • 在心理健康干预等任务中优于基线,提升迁移能力与内容保真度。
  • 适合小样本场景下的复杂语言策略学习,如对话系统设计。

本文提出一种新颖的因果框架,用于自然语言动作空间中的多阶段决策,其中结果仅在一系列动作后才可观察。尽管近期方法如近端策略优化(PPO)可在高维动作空间中处理延迟奖励问题,但通常需要多个模型(策略、价值、奖励)并依赖大量训练数据。本方法采用Q-learning通过单一模型估计动态治疗方案(DTR),利用梯度上升优化语言嵌入,实现数据高效的策略学习。关键技术贡献在于一种解码策略,可将优化后的嵌入转换为连贯自然语言。我们在心理健康干预、仇恨言论应对和情感迁移任务上评估该方法,结果显示在多个指标上显著优于竞争性基线。特别地,该方法在保持内容保留和流畅性的前提下,展现出更强的迁移能力,经人工评估验证。本工作为有限训练数据条件下复杂语言任务的最优策略学习提供了实用基础。

原文摘要 · Abstract (English)

This paper introduces a novel causal framework for multi-stage decision-making in natural language action spaces where outcomes are only observed after a sequence of actions. While recent approaches like Proximal Policy Optimization (PPO) can handle such delayed-reward settings in high-dimensional action spaces, they typically require multiple models (policy, value, and reward) and substantial training data. Our approach employs Q-learning to estimate Dynamic Treatment Regimes (DTR) through a single model, enabling data-efficient policy learning via gradient ascent on language embeddings. A key technical contribution of our approach is a decoding strategy that translates optimized embeddings back into coherent natural language. We evaluate our approach on mental health intervention, hate speech countering, and sentiment transfer tasks, demonstrating significant improvements over competitive baselines across multiple metrics. Notably, our method achieves superior transfer strength while maintaining content preservation and fluency, as validated through human evaluation. Our work provides a practical foundation for learning optimal policies in complex language tasks where training data is limited.

策略学习因果推理自然语言生成小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。