arXiv:2411.10422cs.MAcs.AI2024-11中稿 · ACL被引 4

用游戏测试大模型的创意与骗术能力,发现生僻词影响推理表现。

Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash

  • 构建多智能体博弈框架,让大模型玩虚构定义游戏
  • 实测显示模型欺骗成功率与正确猜解率存在差异
  • 生僻词汇输入导致规则理解与历史记忆能力下降

大型语言模型在复杂任务和交互环境中的表现令人印象深刻,但其创造力仍缺乏深入研究。本文提出一种基于游戏Balderdash的仿真框架,用于评估大模型的创造力与逻辑推理能力。在Balderdash游戏中,玩家需为生僻词编造虚构定义以误导他人,并识别真实定义。本框架使多个大模型智能体参与游戏,通过中央游戏引擎与裁判大模型评估语义等价性。实验分析了不同大模型的表现,考察真定义识别率、欺骗率与正确猜测率等指标。结果揭示:当输入包含不常见词汇时,模型对游戏规则和历史上下文的理解能力显著下降,暴露出推理短板。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive capabilities in complex tasks and interactive environments, yet their creativity remains underexplored. This paper introduces a simulation framework utilizing the game Balderdash to evaluate both the creativity and logical reasoning of LLMs. In Balderdash, players generate fictitious definitions for obscure terms to deceive others while identifying correct definitions. Our framework enables multiple LLM agents to participate in this game, assessing their ability to produce plausible definitions and strategize based on game rules and history. We implemented a centralized game engine featuring various LLMs as participants and a judge LLM to evaluate semantic equivalence. Through a series of experiments, we analyzed the performance of different LLMs, examining metrics such as True Definition Ratio, Deception Ratio, and Correct Guess Ratio. The results provide insights into the creative and deceptive capabilities of LLMs, highlighting their strengths and areas for improvement. Specifically, the study reveals that infrequent vocabulary in LLMs' input leads to poor reasoning on game rules and historical context (https://github.com/ParsaHejabi/Simulation-Framework-for-Multi-Agent-Balderdash).

大模型评估游戏模拟创意生成多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。