arXiv:2509.08785cs.AIcs.MA2025-09

用叙事框架影响AI决策,探索语言模型如何改变强化学习行为

Narrative-Guided Reinforcement Learning: A Platform for Studying Language Model Influence on Decision Making

  • 双系统架构:强化学习决策+语言模型叙事推理
  • 在可配置网格世界中验证叙事对动作选择的影响
  • 适合研究语言模型与决策系统交互的学者

我们提出一个初步实验平台,探究叙事元素如何通过结合强化学习(RL)与语言模型推理来影响人工智能决策。尽管当前AI既具备决策能力又可进行叙事推理,但两者通常被独立研究。本平台采用双系统架构,考察叙事框架对基于奖励的学习的影响。系统包含一个基于过往经验提出动作建议的强化学习策略,以及一个利用不同叙事框架处理这些建议的语言模型,以引导最终决策。该设置在保持环境与奖励结构一致的前提下,实现对叙事要素的初步实验。我们在可配置的网格世界环境中实现该架构,智能体同时接收策略建议和环境信息。模块化设计支持对环境复杂度、叙事参数及强化学习与叙事决策间交互关系的受控测试。日志系统记录从强化学习策略值到语言模型推理过程再到动作选择模式的基本决策指标。尽管仍处初步阶段,该实现为研究不同叙事框架对基于奖励决策的影响,以及优化学习与符号推理在AI系统中的潜在互动,提供了基础框架。

原文摘要 · Abstract (English)

We present a preliminary experimental platform that explores how narrative elements might shape AI decision-making by combining reinforcement learning (RL) with language model reasoning. While AI systems can now both make decisions and engage in narrative reasoning, these capabilities have mostly been studied separately. Our platform attempts to bridge this gap using a dual-system architecture to examine how narrative frameworks could influence reward-based learning. The system comprises a reinforcement learning policy that suggests actions based on past experience, and a language model that processes these suggestions through different narrative frameworks to guide decisions. This setup enables initial experimentation with narrative elements while maintaining consistent environment and reward structures. We implement this architecture in a configurable gridworld environment, where agents receive both policy suggestions and information about their surroundings. The platform's modular design facilitates controlled testing of environmental complexity, narrative parameters, and the interaction between reinforcement learning and narrative-based decisions. Our logging system captures basic decision metrics, from RL policy values to language model reasoning to action selection patterns. While preliminary, this implementation provides a foundation for studying how different narrative frameworks might affect reward-based decisions and exploring potential interactions between optimization-based learning and symbolic reasoning in AI systems.

强化学习语言模型决策机制叙事推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。