arXiv:2512.08937cs.HCcs.AI2025-12被引 3

对比AI与真人对心理健康建议的回应,发现AI表现更优且更温暖。

When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being

  • 用专家评分比较Reddit热门回复与LLM生成建议
  • GPT-4o在有效性、温度等维度全面优于GPT-5
  • 人类建议可微调后媲美AI,适合融合人机协作场景

寻求建议是核心人类行为,互联网已两次重塑这一过程:先是通过论坛和问答社区实现公众意见众包,如今则由大语言模型(LLMs)推动。然而,这些模型在日常心理健康建议中的质量尚不明确。我们开展了两项研究(共210名参与者),让专家评估高票Reddit回复与LLM生成建议。结果显示,LLM整体评分更高,在有效性、温暖度及再次求助意愿上均显著领先。GPT-4o在除奉承性外所有指标上优于GPT-5,表明性能提升未必带来建议质量改善。第二项研究探讨了人类与算法建议的结合方式,发现人类建议经无感润色后可媲美AI生成内容。研究提出面向融合AI、群体智慧与专家监督的建议系统的设计启示。

原文摘要 · Abstract (English)

Seeking advice is a core human behavior that the internet has reinvented twice: first through forums and Q&A communities that crowdsource public guidance, and now through large language models (LLMs). Yet the quality of this LLM advice for everyday well-being scenarios remains unclear. How does it compare, not only against human comments, but against the wisdom of the online crowd? We ran two studies (N=210) in which experts compared top-voted Reddit advice with LLM-generated advice. LLMs ranked significantly higher overall and on effectiveness, warmth, and willingness to seek advice again. GPT-4o beat GPT-5 on all metrics except sycophancy, suggesting that benchmark gains need not improve advice-giving. In Study-2, we examined how human and algorithmic advice could be combined, and found that human advice can be unobtrusively polished to compete with AI-generated comments. We conclude with design implications for advice-giving agents and ecosystems blending AI, crowd input, and expert oversight.

AI建议心理健康人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。