测试大模型在中英文对话中判断关系的能力,发现韩语表现明显更差。
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
- 构建双语对话数据集SCRIPTS,评估模型推断人物关系的能力。
- 英语模型准确率75%-80%,韩语仅58%-69%,且10%-25%预测错误关系。
- 思维链提示对社交推理帮助小,反而可能放大偏见,适合多语言社交研究者。
随着大语言模型在真实交互中的广泛应用,其在人际沟通中的社会推理能力变得至关重要。为此,我们引入了包含1.1千个对话的SCRIPTS数据集,涵盖英、韩双语,来源为电影剧本,并提出了基于该数据集的社会推理任务,评估模型推断对话双方社会关系(如朋友、恋人)的能力。在该任务上评估九个模型,当前模型在英语数据集上准确率约为75%-80%,在韩语数据集上为58%-69%。此外,模型在两种语言中均有10%-25%的响应预测为‘不太可能’的关系。进一步研究发现,思考类模型与思维链提示对社会推理提升有限,甚至可能加剧社会偏见。总体而言,当前大模型在社会推理方面存在显著局限,尤其在韩语中更为明显,凸显了跨语言社交感知型大模型发展的迫切需求。
原文摘要 · Abstract (English)
As LLMs are increasingly deployed in real-world interactions, their social reasoning in interpersonal communication becomes critical. To explore their capabilities, we introduce SCRIPTS, a 1.1k-dialogue dataset in English and Korean, sourced from movie scripts and propose a social reasoning task based on SCRIPTS that evaluates the capacity of LLMs to infer the social relationships (e.g., friends, lovers) between speakers in each dialogue. Evaluating nine models on our task, current LLMs achieve around 75--80% on the English dataset and 58--69% in Korean, and models predict an Unlikely relationship in 10--25% of responses in both languages. Furthermore, we find that thinking models and chain-of-thought prompting provide minimal benefits for social reasoning and occasionally amplify social biases. In sum, there are significant limitations in current LLMs' social reasoning capabilities, especially for Korean, highlighting the need for efforts to develop socially-aware LLMs across languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。