测试大模型能否把人类指令翻译成智能体的内部符号表示
Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
- 用结构化框架评估大模型将自然语言转为智能体符号表征的能力
- 在蚂蚁迷宫和坠落环境中,模型翻译准确率随划分粒度变细而下降
- 揭示当前大模型在语言与内部表示间对齐的局限性,适合研究通用智能体者参考
涌现的符号表征对发展性学习智能体实现规划与跨任务泛化至关重要。本文研究大语言模型(LLMs)是否能将人类自然语言指令转化为分层强化学习中产生的内部符号表示。我们采用结构化评估框架,在蚂蚁迷宫(Ant Maze)和蚂蚁坠落(Ant Fall)环境中,考察GPT、Claude、Deepseek和Grok等常见大模型对不同层级符号划分的翻译性能。结果表明,尽管大模型具备一定将自然语言映射到环境动态符号表征的能力,但其表现高度依赖于划分粒度和任务复杂度。该研究暴露了当前大模型在表示对齐上的局限性,凸显了未来在语言与智能体内部表征间建立鲁棒对齐机制的必要性。
原文摘要 · Abstract (English)
Emergent symbolic representations are critical for enabling developmental learning agents to plan and generalize across tasks. In this work, we investigate whether large language models (LLMs) can translate human natural language instructions into the internal symbolic representations that emerge during hierarchical reinforcement learning. We apply a structured evaluation framework to measure the translation performance of commonly seen LLMs -- GPT, Claude, Deepseek and Grok -- across different internal symbolic partitions generated by a hierarchical reinforcement learning algorithm in the Ant Maze and Ant Fall environments. Our findings reveal that although LLMs demonstrate some ability to translate natural language into a symbolic representation of the environment dynamics, their performance is highly sensitive to partition granularity and task complexity. The results expose limitations in current LLMs capacity for representation alignment, highlighting the need for further research on robust alignment between language and internal agent representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。