arXiv:2607.05956cs.AIcs.CL2026-07

将知识图谱与多语言学术语料融合,提升人文社科领域大模型的适配性与可靠性。

Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

  • 融合知识图谱与多语言文献库,实现领域自适应的大模型训练。
  • 在欧盟LLMs4EU框架下,验证了检索、摘要、溯源与幻觉检测的性能提升。
  • 适合关注数字人文、跨语言研究与生成式AI伦理的学者使用。

大型语言模型(LLMs)融入社会科学与人文学科(SSH)的研究流程,在文献发现与综述方面带来方法论、认识论和监管挑战,尤其涉及学科多样性、多语言资源获取及结果评估。本文介绍欧洲项目LLMs4EU与ALT-EDIC基础设施中的一个正在进行的用例,旨在将基础模型适配至SSH研究实践,支持问答、比较文档分析与文献综述等任务。评估遵循LLMs4EU协议,包含独立的定量基准测试(检索、摘要、可追溯性与幻觉检测)及由数字人文专家组成的定性评估。通过将模型适配嵌入研究基础设施,并在结构化的法律与伦理合规框架内进行,该用例探索了敏感于领域且符合法规的生成式AI如何在保障可靠性和认识责任的前提下支持SSH学术研究。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature synthesis, raises significant methodological, epistemic and regulatory challenges for the Social Sciences and Humanities (SSH), especially with regard to disciplinary diversity, multilingual access to sources and the evaluation of results. This paper presents an on-going use case developed within the European project LLMs4EU and the ALT-EDIC infrastructure, aimed at adapting foundation models to SSH research practices and supporting tasks such as question answering, comparative document analysis and literature review. The evaluation framework follows the LLMs4EU protocol and encompasses both independent quantitative benchmarking (retrieval, summarisation, traceability and hallucination detection) and a qualitative assessment involving a panel of Digital Humanities experts. By embedding model adaptation within research infrastructures and a structured legal and ethical compliance framework, the use case explores how domain-sensitive and regulation-aware generative AI can support SSH scholarship while preserving reliability and epistemic responsibility.

人文社科多语言知识图谱生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。