测试大模型模仿人类联想思维的互动能力,推动更自然的人机协作。
Word Synchronization Challenge: A Benchmark for Word Association Responses for Large Language Models
- 设计游戏化动态框架,通过词语联想评测模型认知模仿能力。
- 模型复杂度越高,越能贴近人类思维模式进行回应。
- 适合研究人机交互、认知建模及情感化智能系统的学者。
本文提出词同步挑战(Word Synchronization Challenge),一项用于评估大语言模型(LLMs)在人机交互(HCI)中表现的新基准。该基准采用动态游戏化框架,通过词语关联任务测试模型模仿人类认知过程的能力。通过模拟复杂的交互情境,评估大模型在对话中理解并匹配人类思维模式的表现,这对实现有效的人机社会协作至关重要。初步结果显示,模型复杂度显著影响其表现,揭示了模型在开展有意义社交互动和以类人方式调整行为方面的潜力。该研究深化了对大模型复制或偏离人类认知功能的理解,为构建更细致、更具同理心的人机合作系统奠定基础。
原文摘要 · Abstract (English)
This paper introduces the Word Synchronization Challenge, a novel benchmark to evaluate large language models (LLMs) in Human-Computer Interaction (HCI). This benchmark uses a dynamic game-like framework to test LLMs ability to mimic human cognitive processes through word associations. By simulating complex human interactions, it assesses how LLMs interpret and align with human thought patterns during conversational exchanges, which are essential for effective social partnerships in HCI. Initial findings highlight the influence of model sophistication on performance, offering insights into the models capabilities to engage in meaningful social interactions and adapt behaviors in human-like ways. This research advances the understanding of LLMs potential to replicate or diverge from human cognitive functions, paving the way for more nuanced and empathetic human-machine collaborations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。