arXiv:2504.21605cs.CLcs.AI2025-04

用RDF结构化表示多语言大模型评估,识别知识冲突与上下文优先级。

RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations

  • 基于RDF构建多语言模型评估的结构化表示框架
  • 在28个问题中覆盖全部评估维度,发现语言差异与上下文依赖模式
  • 适合关注模型可靠性、跨语言一致性与事实偏倚的研究者

大语言模型日益作为知识接口,但系统评估其在存在矛盾信息时的可靠性仍具挑战。我们提出一种基于RDF的框架,用于多语言大模型质量评估,聚焦知识冲突。该方法在德语和英语下,针对四种不同上下文条件(完整、不完整、冲突、无上下文)捕捉模型响应。此结构化表示支持对知识泄露(模型偏好训练数据而非提供上下文)、错误检测及多语言一致性进行全面分析。通过一个消防安全领域的实验验证,揭示了上下文优先级的关键模式与语言特异性表现,并证明所用词汇足以表达28个问题中的所有评估维度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly serve as knowledge interfaces, yet systematically assessing their reliability with conflicting information remains difficult. We propose an RDF-based framework to assess multilingual LLM quality, focusing on knowledge conflicts. Our approach captures model responses across four distinct context conditions (complete, incomplete, conflicting, and no-context information) in German and English. This structured representation enables the comprehensive analysis of knowledge leakage-where models favor training data over provided context-error detection, and multilingual consistency. We demonstrate the framework through a fire safety domain experiment, revealing critical patterns in context prioritization and language-specific performance, and demonstrating that our vocabulary was sufficient to express every assessment facet encountered in the 28-question study.

大模型评估RDF多语言知识冲突

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。