构建中文微博道德语料库,助力自然语言中道德情感分析
The Moral Foundations Weibo Corpus
- 基于十类道德理论手工标注2.5万条微博评论
- 三名标注员一致性达0.78以上(Kappa值)
- 适配中文NLP研究者与伦理计算方向应用
自然语言中的道德情感显著影响线上线下环境,塑造行为模式与互动方式,涵盖社交媒体自我呈现、网络欺凌、社会规范遵从及伦理决策。为有效测量自然语言中的道德情感,需依赖大规模、精细化标注的数据集以支持准确分析与模型训练。然而现有语料库在汉语领域存在语言局限性。为此,本文提出「道德基础微博语料库」,包含25,671条来自微博的中文评论,覆盖六个不同话题领域。每条评论由至少三位系统培训的标注员基于源于道德根基理论的十类道德维度进行人工标注。通过卡帕系数检验标注者一致性,结果表明标注可靠。此外,引入多款前沿大模型辅助标注,开展对比实验并报告道德情感分类基线性能。
原文摘要 · Abstract (English)
Moral sentiments expressed in natural language significantly influence both online and offline environments, shaping behavioral styles and interaction patterns, including social media selfpresentation, cyberbullying, adherence to social norms, and ethical decision-making. To effectively measure moral sentiments in natural language processing texts, it is crucial to utilize large, annotated datasets that provide nuanced understanding for accurate analysis and modeltraining. However, existing corpora, while valuable, often face linguistic limitations. To address this gap in the Chinese language domain,we introduce the Moral Foundation Weibo Corpus. This corpus consists of 25,671 Chinese comments on Weibo, encompassing six diverse topic areas. Each comment is manually annotated by at least three systematically trained annotators based on ten moral categories derived from a grounded theory of morality. To assess annotator reliability, we present the kappa testresults, a gold standard for measuring consistency. Additionally, we apply several the latest large language models to supplement the manual annotations, conducting analytical experiments to compare their performance and report baseline results for moral sentiment classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。