arXiv:2512.12443cs.AIcs.SE2025-12被引 2

构建AI模型透明度评估框架,自动打分并发现安全披露短板

AI Transparency Atlas: Framework, Scoring, and Real-Time Model Card Evaluation Pipeline

  • 设计8大模块23小节的加权透明度框架,突出安全评估权重
  • 自动化评估50个模型仅耗时3美元,发现主流厂商合规率不足60%
  • 揭露前沿模型在幻觉、欺骗行为等关键安全项上普遍缺失

AI模型文档分散于各平台且结构不一,阻碍政策制定者、审计人员和用户可靠评估安全性、数据来源及版本变更。我们分析了五个前沿模型(Gemini 3、Grok 4.1、Llama 4、GPT-5、Claude 4.5)和100个Hugging Face模型卡,识别出947个不同的章节名称,使用信息仅以97种不同标签出现。基于欧盟《人工智能法案》附录IV与斯坦福透明度指数,构建包含8个部分、23个子部分的加权透明度框架,其中安全评估占25%,关键风险占20%。开发自动化多智能体管道,从公开源提取文档并通过大模型共识评分。对50个视觉、多模态、开源与闭源模型的评估总成本低于3美元,揭示系统性缺失:前沿实验室(xAI、微软、Anthropic)合规率约80%,多数提供方低于60%。安全关键类别缺口最大:欺骗行为、幻觉、儿童安全评估分别累计丢失148、124、116分。

原文摘要 · Abstract (English)

AI model documentation is fragmented across platforms and inconsistent in structure, preventing policymakers, auditors, and users from reliably assessing safety claims, data provenance, and version-level changes. We analyzed documentation from five frontier models (Gemini 3, Grok 4.1, Llama 4, GPT-5, and Claude 4.5) and 100 Hugging Face model cards, identifying 947 unique section names with extreme naming variation. Usage information alone appeared under 97 distinct labels. Using the EU AI Act Annex IV and the Stanford Transparency Index as baselines, we developed a weighted transparency framework with 8 sections and 23 subsections that prioritizes safety-critical disclosures (Safety Evaluation: 25%, Critical Risk: 20%) over technical specifications. We implemented an automated multi-agent pipeline that extracts documentation from public sources and scores completeness through LLM-based consensus. Evaluating 50 models across vision, multimodal, open-source, and closed-source systems cost less than $3 in total and revealed systematic gaps. Frontier labs (xAI, Microsoft, Anthropic) achieve approximately 80% compliance, while most providers fall below 60%. Safety-critical categories show the largest deficits: deception behaviors, hallucinations, and child safety evaluations account for 148, 124, and 116 aggregate points lost, respectively, across all evaluated models.

模型透明度安全评估自动化评测AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。