arXiv:2511.15097cs.CRcs.AI2025-11

用可验证的数据构件构建可信AI,让每步操作都可审计。

MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm

  • 以持久数据构件驱动AI行为,替代临时任务
  • 实现2720.7 MB/s高速流处理与225倍压缩率
  • 适合需要合规、安全与可追溯的高风险场景

人工智能信任危机正阻碍其在关键领域的应用,监管障碍、安全漏洞和责任缺失成为主要瓶颈。现有AI系统依赖不透明的数据结构,缺乏审计轨迹、来源追踪和可解释性,难以满足欧盟《人工智能法案》等新规要求。本文提出一种以数据构件为中心的AI代理范式,通过持久且可验证的数据构件驱动行为,从数据架构层面解决信任问题。核心是多模态数据构件文件格式(MAIF),它嵌入语义表示、密码学溯源和细粒度访问控制,使数据从被动存储变为主动信任保障,确保每项AI操作均可审计。生产级实现达到2720.7 MB/s超高速流传输、1342 MB/s优化视频处理能力,支持最高225倍压缩比的同时保持语义保真。新增跨模态注意力、语义压缩与密码绑定算法,具备流级别访问控制、实时篡改检测与行为异常分析能力,开销极低。该方案直接应对监管、安全与问责挑战,为大规模可信AI部署提供可行路径。

原文摘要 · Abstract (English)

The AI trustworthiness crisis threatens to derail the artificial intelligence revolution, with regulatory barriers, security vulnerabilities, and accountability gaps preventing deployment in critical domains. Current AI systems operate on opaque data structures that lack the audit trails, provenance tracking, or explainability required by emerging regulations like the EU AI Act. We propose an artifact-centric AI agent paradigm where behavior is driven by persistent, verifiable data artifacts rather than ephemeral tasks, solving the trustworthiness problem at the data architecture level. Central to this approach is the Multimodal Artifact File Format (MAIF), an AI-native container embedding semantic representations, cryptographic provenance, and granular access controls. MAIF transforms data from passive storage into active trust enforcement, making every AI operation inherently auditable. Our production-ready implementation demonstrates ultra-high-speed streaming (2,720.7 MB/s), optimized video processing (1,342 MB/s), and enterprise-grade security. Novel algorithms for cross-modal attention, semantic compression, and cryptographic binding achieve up to 225 compression while maintaining semantic fidelity. Advanced security features include stream-level access control, real-time tamper detection, and behavioral anomaly analysis with minimal overhead. This approach directly addresses the regulatory, security, and accountability challenges preventing AI deployment in sensitive domains, offering a viable path toward trustworthy AI systems at scale.

可信AI数据溯源安全框架多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。