研究发现大模型像人一样,适度压力下表现最好。
StressPrompt: Does Stress Impact Large Language Models and Human Performance Similarly?
- 设计了模拟人类压力的新型提示词StressPrompt。
- 大模型在中等压力下表现最优,高低压力均导致性能下降。
- 提示词改变模型内部表征,与人类应激反应相似,适合关注AI稳定性者。
人类常经历压力,显著影响表现。本研究探讨大语言模型(LLMs)是否具有类似人类的压力反应,以及其在不同压力提示下的表现波动。我们开发了一种名为StressPrompt的新颖提示集,基于心理学框架并经由人类参与者评分校准,以诱发不同程度的压力。将这些提示应用于多个LLMs,评估其在指令遵循、复杂推理和情感智能任务中的表现。结果表明,LLMs的表现也遵循耶克斯-多德森定律,在中等压力下达到最佳,低或高压下均下降。分析显示,这些提示显著改变了模型的内部状态,使其神经表示变化与人类压力反应相仿。该研究揭示了大模型在真实场景中(如客服、医疗、应急响应)维持高性能的重要性,为提升系统鲁棒性提供依据,并为理解大模型认知机制提供了新视角。
原文摘要 · Abstract (English)
Human beings often experience stress, which can significantly influence their performance. This study explores whether Large Language Models (LLMs) exhibit stress responses similar to those of humans and whether their performance fluctuates under different stress-inducing prompts. To investigate this, we developed a novel set of prompts, termed StressPrompt, designed to induce varying levels of stress. These prompts were derived from established psychological frameworks and carefully calibrated based on ratings from human participants. We then applied these prompts to several LLMs to assess their responses across a range of tasks, including instruction-following, complex reasoning, and emotional intelligence. The findings suggest that LLMs, like humans, perform optimally under moderate stress, consistent with the Yerkes-Dodson law. Notably, their performance declines under both low and high-stress conditions. Our analysis further revealed that these StressPrompts significantly alter the internal states of LLMs, leading to changes in their neural representations that mirror human responses to stress. This research provides critical insights into the operational robustness and flexibility of LLMs, demonstrating the importance of designing AI systems capable of maintaining high performance in real-world scenarios where stress is prevalent, such as in customer service, healthcare, and emergency response contexts. Moreover, this study contributes to the broader AI research community by offering a new perspective on how LLMs handle different scenarios and their similarities to human cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。