arXiv:2510.25908cs.AI2025-10

评测大模型在科研中的可信度,覆盖真实、安全、伦理等四维度。

SciTrust 2.0: A Comprehensive Framework for Evaluating Trustworthiness of Large Language Models in Scientific Applications

  • 构建四维评估框架,含真实、抗攻击、安全、伦理四个维度。
  • 通用模型整体优于专业模型,GPT-o4-mini在真实性和抗攻击上表现最佳。
  • 科学专用模型存在逻辑与伦理短板,高风险领域安全漏洞明显。

大型语言模型(LLMs)在科研中展现出变革潜力,但在高风险场景下的可信度仍存疑。本文提出SciTrust 2.0,一个涵盖真实性、对抗鲁棒性、科学安全性和科学伦理的综合性评估框架。通过验证的反思调优流程和专家验证,构建了开放式真实性基准,并开发了覆盖双用途研究、偏见等八个子类别的科学伦理基准。评估了七款主流LLM,包括四款科学专用模型和三款通用行业模型,采用准确率、语义相似度及基于LLM的评分等多指标。结果显示,通用模型在各维度上整体优于科学专用模型,GPT-o4-mini在真实性与对抗鲁棒性上表现最优;而科学专用模型在逻辑与伦理推理方面存在明显缺陷,且在生物安全、化学武器等高风险领域存在显著安全隐患。本框架已开源,为构建更可信的AI系统提供基础,推动科研场景下模型安全与伦理研究。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated transformative potential in scientific research, yet their deployment in high-stakes contexts raises significant trustworthiness concerns. Here, we introduce SciTrust 2.0, a comprehensive framework for evaluating LLM trustworthiness in scientific applications across four dimensions: truthfulness, adversarial robustness, scientific safety, and scientific ethics. Our framework incorporates novel, open-ended truthfulness benchmarks developed through a verified reflection-tuning pipeline and expert validation, alongside a novel ethics benchmark for scientific research contexts covering eight subcategories including dual-use research and bias. We evaluated seven prominent LLMs, including four science-specialized models and three general-purpose industry models, using multiple evaluation metrics including accuracy, semantic similarity measures, and LLM-based scoring. General-purpose industry models overall outperformed science-specialized models across each trustworthiness dimension, with GPT-o4-mini demonstrating superior performance in truthfulness assessments and adversarial robustness. Science-specialized models showed significant deficiencies in logical and ethical reasoning capabilities, along with concerning vulnerabilities in safety evaluations, particularly in high-risk domains such as biosecurity and chemical weapons. By open-sourcing our framework, we provide a foundation for developing more trustworthy AI systems and advancing research on model safety and ethics in scientific contexts.

大模型评估科学可信度伦理安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。