arXiv:2601.09017cs.CL2026-01被引 1

用多语言社交推理游戏评估大模型的文化适应能力。

Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game

  • 设计动态多语言社交推理游戏,测试模型在文化情境下的对话策略。
  • 非英语语境下模型表现显著下降,尤其在本地化实体理解和规则遵守上。
  • 提供可扩展、防数据泄露、注重文化细节的新型评估方式,适合评测多语种模型。

大语言模型的快速发展迫切需要超越静态基准的更强大评估方法,而传统基准正面临数据饱和与泄露问题。本文提出一种基于社交推理游戏Spyfall的动态评估框架,用于衡量模型在多语言和跨文化场景下的能力。模型需通过策略性对话识别秘密特工或避免暴露身份,涉及具有文化相关性的地点或地方食物。实验结果显示,该方法的游戏排名与Chatbot Arena高度一致。然而,在非英语语境中存在显著性能差距:模型在处理本地特定实体时普遍表现较差,且在非英语语言中常出现规则遵守或策略一致性问题。本研究证明,这种基于游戏的方法可提供一种可扩展、抗数据泄露且具备文化细腻度的替代方案,适用于传统NLP基准之外的评估需求。游戏历史数据可在HuggingFace仓库https://huggingface.co/datasets/haryoaw/cultural-spyfall获取。

原文摘要 · Abstract (English)

The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturation and leakage. In this paper, we propose a dynamic benchmarking framework for evaluating multilingual and multicultural capabilities through the social deduction game Spyfall. In our setup, models must engage in strategic dialogue to either identify a secret agent or avoid detection, utilizing culturally relevant locations or local foods. Our results show that our game-based rankings align closely with the Chatbot Arena. However, we find a significant performance gap in non-English contexts: models are generally less proficient when handling locally specific entities and often struggle with rule-following or strategic integrity in non-English languages. We demonstrate that this game-based approach provides a scalable, leakage-resistant, and culturally nuanced alternative to traditional NLP benchmarks. The game history can be accessed here https://huggingface.co/datasets/haryoaw/cultural-spyfall.

多语言评估社交推理文化适应大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。