用轻量模型精准识别网络言论中的道德立场,提升跨领域检测效果。
Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour
- 基于心理理论设计三模块架构,融合道德依据与仇恨言论信号。
- 在多个数据集上比现有方法提升15.3%的AUC,跨域性能更优。
- 适合研究虚假信息传播与网络道德行为的学者与平台方使用。
道德语言在塑造网络支持行为和信息扩散中起核心作用,但现有道德基础检测系统普遍存在跨领域泛化差、理由支撑弱及依赖高成本提示型大模型的问题。本文提出CHARM框架,基于轻量微调的大模型,集成互补的道德依据、理由对齐与极性感知的仇恨言论信号,实现更鲁棒、更可信的道德预测。不同于以往字典法、微调法或提示法检测器(这些方法将计算与心理学理论割裂),CHARM每个组件——MAC交叉注意力、理由对齐、仇恨言论调制——均对应一个明确的心理学概念。利用MFTC、MFRC和News训练集的30%子样本,结合更丰富的MFTCXplain监督信号,CHARM在域内提升最高15.3%的AUC,且在所有域外数据集上,于AUC与F1指标均超越有监督基线。该框架可扩展、成本低,是提示型大模型检测的可行替代方案。进一步应用于推特上的大规模新冠疫情话语分析,发现道德价值对齐显著关联线上支持行为。通过实现大规模道德框架可测量,CHARM为研究带有道德色彩的虚假信息传播提供了实用工具。
原文摘要 · Abstract (English)
Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MA\textbf{C}- and \textbf{H}ate-speech-\textbf{A}ware \textbf{R}ationale-aligned \textbf{M}oral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。