用AI推理长度预测人工审核耗时,发现两者高度相关。
AI reasoning effort predicts human decision time in content moderation
- 用链式思维长度衡量AI推理努力,预测人工决策时间。
- 三款前沿模型中,推理越长,人类耗时越久,任务难度影响一致。
- 模型在复杂情境下更频繁引用上下文因素,具可解释性潜力。
大型语言模型可在生成答案前生成中间推理步骤,通过交互式求解提升难题表现。本研究以内容审核任务为场景,考察人类决策时间与模型推理努力(以链式思维,CoT长度衡量)之间的关联。在三款前沿模型中,CoT长度均能稳定预测人类决策时间。此外,当关键变量保持不变时,人类耗时更长,模型生成的CoT也更长,表明对任务难度具有相似敏感性。对CoT内容的分析显示,模型在决策时更频繁引用各类上下文因素。这些发现揭示了人类与AI在实际任务中的推理平行性,强调了推理轨迹在提升可解释性与决策支持方面的潜力。
原文摘要 · Abstract (English)
Large language models can now generate intermediate reasoning steps before producing answers, improving performance on difficult problems by interactively developing solutions. This study uses a content moderation task to examine parallels between human decision times and model reasoning effort, measured using the length of the chain-of-thought (CoT). Across three frontier models, CoT length consistently predicts human decision time. Moreover, humans took longer and models produced longer CoTs when important variables were held constant, suggesting similar sensitivity to task difficulty. Analyses of the CoT content shows that models reference various contextual factors more frequently when making such decisions. These findings show parallels between human and AI reasoning on practical tasks and underscore the potential of reasoning traces for enhancing interpretability and decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。