复杂任务需推理,简单任务反被拖累,实证揭示大模型推理的适用边界。
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
- 对比504种配置,发现推理效果随任务复杂度变化显著
- 二分类任务性能最高下降19.9 F1,27类情绪识别提升16.0 F1
- 仅在复杂情绪识别中值得投入额外计算开销
大型语言模型(LLMs)的推理能力被广泛认为能普遍提升任务表现。我们通过在七种模型族、共504种配置上对不同粒度的情感分析数据集(二分类、五分类、27类情绪)进行综合评估,检验该说法。结果表明:(1)推理有效性具有强任务依赖性——二分类任务性能最高下降19.9 F1百分点,而27类情绪识别最高提升16.0 F1;(2)蒸馏推理版本在简单任务上比基础模型低3-18 F1;(3)少样本提示在多数情况下优于零样本,增益随架构和任务复杂度变化;(4)帕累托前沿分析显示,基础模型在效率-性能权衡中占优,推理仅在复杂情绪识别中合理,尽管带来2.1倍至54倍的计算开销。定性错误分析揭示,推理在简单任务中因系统性过度思辨而退化,提供超越‘过度思考’假说的机制洞察。
原文摘要 · Abstract (English)
Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations across seven model families--including adaptive, conditional, and reinforcement learning-based reasoning architectures--on sentiment analysis datasets of varying granularity (binary, five-class, and 27-class emotion). Our findings reveal that reasoning effectiveness is strongly task-dependent, challenging prevailing assumptions: (1) Reasoning shows task-complexity dependence--binary classification degrades up to -19.9 F1 percentage points (pp), while 27-class emotion recognition gains up to +16.0pp; (2) Distilled reasoning variants underperform base models by 3-18 pp on simpler tasks, though few-shot prompting enables partial recovery; (3) Few-shot learning improves over zero-shot in most cases regardless of model type, with gains varying by architecture and task complexity; (4) Pareto frontier analysis shows base models dominate efficiency-performance trade-offs, with reasoning justified only for complex emotion recognition despite 2.1x-54x computational overhead. We complement these quantitative findings with qualitative error analysis revealing that reasoning degrades simpler tasks through systematic over-deliberation, offering mechanistic insight beyond the high-level overthinking hypothesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。