arXiv:2606.14748cs.CVcs.AI2026-06被引 9

检测AI训练数据是否包含特定样本,助力模型透明与合规。

Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2

论文配图:Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2
图 1 · 摘自论文原文
  • 提出MINT框架,通过实验判断数据是否被用于训练。
  • 在人脸与大模型上实现最高90%的训练数据检测准确率。
  • 开源平台支持图文多模态审计,适合监管与研究使用。

我们提出会员推断测试(MINT)演示2,一个旨在提升机器学习训练过程透明度的框架。MINT是一种实验性技术,可确定特定数据是否被用于模型训练。我们建立了理论框架,并根据对被审计模型信息的掌握程度,提出了多种MINT架构。基于主流人脸识别模型、4个前沿大语言模型及多个多样化、大规模公开图文数据库的实验结果显示,训练数据检测准确率最高达90%。在此基础上,我们推出综合性网络平台1,扩展了对图像与文本模态的审计能力。该平台集成MINT、aMINT与gMINT多种技术,支持多种模型的审计。本演示旨在推动AI透明化,提供符合新兴AI法规要求的实用工具。

原文摘要 · Abstract (English)

We present the Membership Inference Test (MINT) Demo 2, a framework designed to improve transparency in machine learning training processes. MINT is a technique for experimentally determining whether specific data were used during machine learning model training. We establish the theoretical framework and propose multiple architectures for MINT depending on the amount of information known about the models that are being audited. Experimental results using a popular face recognition model, 4 state-of-the-art LLMs, and multiple, diverse, and large-scale public image and text databases achieve promising accuracy levels in the detection of training data of up to 90%. Building on these results, we introduce a comprehensive web platform1 that expands these capabilities to image and text modalities. The platform integrates a diverse technological stack, including MINT, aMINT, and gMINT, allowing users to audit a wide range of models. This demonstrator aims to promote AI transparency and provides a practical tool to foster compliance with emerging AI regulations.

AI透明数据审计模型安全大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。