大模型在道德推理上表现不佳,远不如专门微调的模型。
The Moral Gap of Large Language Models
- 用社交媒体数据对比大模型与微调模型的道德判断能力
- 大模型误检率高,常漏掉道德内容,即使优化提示词也无改善
- 适合关注伦理对齐、道德识别的研究者和开发者参考
道德基础检测对于分析社会话语和开发伦理对齐的AI系统至关重要。尽管大语言模型在多种任务中表现出色,但其在专业道德推理任务上的表现尚不明确。本研究首次基于推特和Reddit数据集,通过ROC、PR和DET曲线分析,全面比较了当前最先进的大语言模型与微调的Transformer模型。结果揭示出显著性能差距:大模型存在高误报率和系统性漏检,即便经过提示工程优化仍无法有效识别道德内容。这些发现表明,在道德推理应用中,任务特定的微调仍优于提示工程。
原文摘要 · Abstract (English)
Moral foundation detection is crucial for analyzing social discourse and developing ethically-aligned AI systems. While large language models excel across diverse tasks, their performance on specialized moral reasoning remains unclear. This study provides the first comprehensive comparison between state-of-the-art LLMs and fine-tuned transformers across Twitter and Reddit datasets using ROC, PR, and DET curve analysis. Results reveal substantial performance gaps, with LLMs exhibiting high false negative rates and systematic under-detection of moral content despite prompt engineering efforts. These findings demonstrate that task-specific fine-tuning remains superior to prompting for moral reasoning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。