用中文验证非英语道德基础测量,发现大模型表现最佳。
Beyond English: Evaluating Automated Measurement of Moral Foundations in Non-English Discourse with a Chinese Case Study
- 对比机器翻译、本地词典、多语言模型和大模型在中文中的应用
- 大模型在跨语言道德判断上表现最好,数据效率高
- 强调人工校验必要性,避免文化细节丢失
本研究探讨了在非英语语料中计算衡量道德基础(MFs)的方法。由于多数资源主要针对英语开发,道德基础理论的跨语言应用仍受限。以中文为例,论文评估了将英语资源应用于机器翻译文本、本地语言词典、多语言语言模型以及大语言模型(LLMs)在非英语文本中测量道德基础的有效性。结果表明,机器翻译和本地词典方法在复杂道德判断中表现不足,常导致文化信息严重丢失。相比之下,多语言模型和大语言模型通过迁移学习展现出可靠的跨语言性能,其中大语言模型在数据效率方面尤为突出。重要的是,研究强调了自动化道德基础评估需人工介入验证,因为最先进模型可能忽略跨语言测量中的文化细微差别。研究结果凸显了大语言模型在跨语言道德基础测量及其他复杂多语言推断编码任务中的潜力。
原文摘要 · Abstract (English)
This study explores computational approaches for measuring moral foundations (MFs) in non-English corpora. Since most resources are developed primarily for English, cross-linguistic applications of moral foundation theory remain limited. Using Chinese as a case study, this paper evaluates the effectiveness of applying English resources to machine translated text, local language lexicons, multilingual language models, and large language models (LLMs) in measuring MFs in non-English texts. The results indicate that machine translation and local lexicon approaches are insufficient for complex moral assessments, frequently resulting in a substantial loss of cultural information. In contrast, multilingual models and LLMs demonstrate reliable cross-language performance with transfer learning, with LLMs excelling in terms of data efficiency. Importantly, this study also underscores the need for human-in-the-loop validation of automated MF assessment, as the most advanced models may overlook cultural nuances in cross-language measurements. The findings highlight the potential of LLMs for cross-language MF measurements and other complex multilingual deductive coding tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。