为多语言大模型在人文学科中提供可解释的评估框架
Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities
- 将多语言大模型视为意义生成的诠释工具,而非单纯任务完成器
- 提出文化契合度、跨语言稳定性等可量化评估指标
- 适合关注伦理与跨文化研究的计算人文学家使用
大语言模型在多语言能力和推理方面快速进步,已可融入社会科学与人文学科的研究流程。然而现有评估仍局限于任务型NLP基准,未能回应解释有效性、文化语境嵌入性与认识论中介等问题。本文将多语言推理大模型重新定义为在语言与文化语境中主动建构意义的诠释工具,结合诠释学、科技哲学、多语言NLP与计算社会科学方法,构建一个理论扎实的评估框架。提出包含文化契合度、跨语言稳定性与推理忠实性在内的可操作化指标体系,并针对解释性研究任务设定透明性要求。通过多语言政治话语分析的应用案例展示该框架。论文为人文学科中负责任地整合多语言推理大模型提供了概念与方法基础。
原文摘要 · Abstract (English)
Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humanities research workflows. Yet existing evaluation paradigms remain anchored in task-based NLP benchmarks and fail to address interpretive validity, cultural situatedness, and epistemic mediation. This paper reconceptualizes multilingual reasoning LLMs as hermeneutic instruments that actively structure meaning production across linguistic and cultural contexts. Drawing on hermeneutics, philosophy of technology, science and technology studies, multilingual NLP research, and computational social science methodology, we develop a theoretically grounded framework for evaluating multilingual reasoning in Social Sciences and Humanities (SSH) research. We articulate a rigorous experimental protocol with operationalized metrics for cultural alignment, cross-lingual stability, and reasoning faithfulness, along with transparency requirements tailored to interpretive research tasks. We illustrate the framework through a concrete application scenario involving multilingual political discourse analysis. The paper contributes a conceptual and methodological foundation for responsible integration of multilingual reasoning LLMs into computational social science infrastructures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。