用语言激励提升大模型自信心和表现,效果因模型而异。
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
- 设计三类言语激励提示:鼓励、挑衅、批评,激发模型自信。
- 不同激励对多数任务有效,最佳类型随模型变化。
- 实验验证心理理论,为模型行为研究提供新视角。
大型语言模型(LLMs)的零样本能力显著提升。由于对输入高度敏感,研究更关注通过简单提示工程而非复杂领域适配来增强性能。研究表明,LLMs具备情感智能,正负情绪均可能提升任务表现。但以往提示多仅使用单一刺激类型,未比较不同刺激效果、未考察任务难度差异影响,也缺乏机制探究。本文受社会认知理论中自我效能与任务表现正相关启发,提出言语自我效能激励(Verbal Efficacy Stimulations, VES),包含鼓励、挑衅、批评三类提示,覆盖帮助性、能力等六个维度,并按难度分类任务,系统研究不同VES对模型自我效能与任务表现的影响。实验表明,三类VES普遍提升模型性能,最优激励类型因模型而异。大量实验结果与心理学理论一致,为未来研究提供了新洞见。
原文摘要 · Abstract (English)
Significant improvements have been observed in the zero-shot capabilities of the Large Language Models (LLMs). Due to their high sensitivity to input, research has increasingly focused on enhancing LLMs' performance via direct and simple prompt engineering rather than intricate domain adaptation. Studies suggest that LLMs exhibit emotional intelligence, and both positive and negative emotions can potentially enhance task performances. However, prior interaction prompts have predominantly concentrated on a single stimulus type, neglecting to compare different stimulus effects, examine the influence of varying task difficulties, or explore underlying mechanisms. This paper, inspired by the positive correlation between self-efficacy and task performance within the social cognitive theory, introduces Verbal Efficacy Stimulations (VES). Our VES comprises three types of verbal prompts: encouraging, provocative, and critical, addressing six aspects such as helpfulness and competence. And we further categorize task difficulty, aiming to extensively investigate how distinct VES influence the self-efficacy and task achievements of language models at varied levels of difficulty. The experimental results show that the three types of VES improve the performance of LLMs on most tasks, and the most effective VES varies for different models. In extensive experiments, we have obtained some findings consistent with psychological theories, providing novel insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。