用情绪调控让大模型跳出思维死胡同,提升解题准确率。
HEART: Emotionally-Driven Test-Time Scaling of Language Models
- 通过交替使用批判与鼓励语气,引导模型聚焦关键问题
- 在7个高难度基准上均超越无情感调节的基线模型
- 适合需要深度推理的复杂任务,如数学、编程与知识问答
测试时扩展显著提升了人工智能模型的问题求解能力,但现有方法常陷入重复错误的思维模式。我们提出HEART框架,利用情感线索引导模型注意力,如同情绪在人类决策中的作用。通过交替采用批判性语气以强化错误识别,以及鼓励性语气以激发新思路,HEART帮助模型突破思维僵局,找到正确解法。我们在七个高难度基准(包括Humanity's Last Exam、GPQA Diamond和LiveCodeBench)上评估了HEART,验证了其在多种模型上的鲁棒性。结果表明,情感调节能促进更深层次的推理,在多个任务中实现稳定准确率提升。研究提示,机器推理的未来在于有策略地融合情感调控以引导逻辑整合。
原文摘要 · Abstract (English)
Test-time scaling has significantly improved how AI models solve problems, yet current methods often get stuck in repetitive, incorrect patterns of thought. We introduce HEART, a framework that uses emotional cues to guide the model's focus, much like how feelings contribute to human decision-making. By alternating between critical tones to sharpen error detection and encouraging tones to spark new ideas, HEART helps the model break out of dead-end reasoning and find the right solution. We evaluate HEART across seven high-difficulty benchmarks--including Humanity's Last Exam, GPQA Diamond, and LiveCodeBench--demonstrating robustness across diverse models. Results show that emotion facilitates deeper reasoning, yielding consistent accuracy gains over affect-sterile baselines. These findings suggest that the next frontier in machine reasoning lies in the strategic integration of affective regulation to guide logical synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。