arXiv:2502.06470cs.CLcs.AI2025-02综述被引 10
梳理大模型心智理论能力,揭示其安全风险与评估方向
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks
- 系统评估大模型在行为与表征层面的心智理论能力
- 发现高阶心智理论可能引发新型安全风险
- 提出可解释性评估与风险缓解的研究路径
心智理论(ToM)是理解他人心理状态并预测其行为的社会智能基础。本文综述了评估大语言模型(LLMs)在行为与表征层面心智理论能力的研究,识别出高级心智理论能力带来的潜在安全风险,并提出了若干有效评估与缓解这些风险的研究方向。
原文摘要 · Abstract (English)
Theory of Mind (ToM), the ability to attribute mental states to others and predict their behaviour, is fundamental to social intelligence. In this paper, we survey studies evaluating behavioural and representational ToM in Large Language Models (LLMs), identify important safety risks from advanced LLM ToM capabilities, and suggest several research directions for effective evaluation and mitigation of these risks.
心智理论大模型安全评估方法
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。