arXiv:2509.07389cs.CLcs.AI2025-09被引 1

测试大模型能否通过对话互动学会新语言,发现其学习策略类似人类。

Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents

  • 让大模型与只懂新语言的机器人对话,观察其语言习得能力。
  • 100次回复内无法建立有效对话,但展现出类人学习策略。
  • 为评估框架和模型设计提供新方向,适合关注交互学习的研究者。

现有大语言模型(LLM)语言能力评估主要聚焦词汇、形态规则、句法泛化、语用推理和跨语言迁移,但均未考察模型是否能通过模式识别和交互反馈习得语言——这是人类语言习得的核心特征。本文提出一种新实验框架:让一个LLM代理与仅理解新构造语言Tinkatongue的机器人进行对话,评估其语言习得与应用能力。结果表明,该代理在100次响应内未能建立有效对话,但采用了多种策略,其行为模式与人类语言学习方式相似。研究揭示了当前评估范式的局限性,提出了新的基准方向,并为能够更高效从交互反馈中学习的模型设计开辟路径。

原文摘要 · Abstract (English)

Existing evaluation studies on linguistic competence of large language models (LLM agents) have focused primarily on vocabulary learning, morphological rule induction, syntactic generalization, pragmatic inference, and cross-linguistic transfer. However, none assess whether LLM agents can acquire a language through pattern recognition and interactive feedback, a central feature of human language acquisition. We propose a novel experimental framework in which an LLM agent is evaluated on its ability to acquire and use a newly constructed language (Tinkatongue) in conversation with a bot that understands only Tinkatongue. Our findings show that LLM agents fail to establish a conversation within 100 responses, yet they adopt distinct strategies that mirror human approaches to language learning. The results suggest a new direction for evaluation benchmarks and open pathways to model designs that learn more effectively from interactive feedback.

语言模型交互学习评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。