arXiv:2603.08639cs.CVcs.AI2026-03

无需模型内部信息,用自然语言揭示视觉模型的隐含认知。

UNBOX: Unveiling Black-box visual models with Natural-language

  • 用大语言模型和文生图模型将激活最大化转为语义搜索。
  • 在ImageNet-1K等数据集上生成高保真语义描述,性能媲美白盒方法。
  • 适合需审计黑箱模型、检测偏见或分析失败案例的研究者。

开放世界视觉识别的可信性依赖于可解释、公平且对分布偏移鲁棒的模型。然而现代视觉系统越来越多以专有黑箱API形式部署,仅暴露输出概率,隐藏架构、参数、梯度与训练数据。这种不透明性阻碍了有意义的审计、偏见检测与故障分析。现有解释方法依赖白盒或灰盒访问,或假设已知训练分布,在真实场景中不可用。我们提出UNBOX框架,实现完全无数据、无梯度、无反向传播条件下的类别级模型剖析。UNBOX利用大语言模型与文本到图像扩散模型,将激活最大化重构为由输出概率驱动的纯语义搜索。该方法生成能最大激活各分类的人类可读文本描述,揭示模型隐含学习的概念、反映的训练分布及潜在偏见来源。我们在ImageNet-1K、Waterbirds和CelebA上通过语义保真度测试、视觉特征相关性分析与切片发现审计进行评估。尽管在最严格的黑箱约束下运行,UNBOX仍表现优于现有方法,证明在无内部访问条件下亦可恢复对模型内在推理的有意义洞察,推动更可信、可问责的视觉识别系统发展。

原文摘要 · Abstract (English)

Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet modern vision systems are increasingly deployed as proprietary black-box APIs, exposing only output probabilities and hiding architecture, parameters, gradients, and training data. This opacity prevents meaningful auditing, bias detection, and failure analysis. Existing explanation methods assume white- or gray-box access or knowledge of the training distribution, making them unusable in these real-world settings. We introduce UNBOX, a framework for class-wise model dissection under fully data-free, gradient-free, and backpropagation-free constraints. UNBOX leverages Large Language Models and text-to-image diffusion models to recast activation maximization as a purely semantic search driven by output probabilities. The method produces human-interpretable text descriptors that maximally activate each class, revealing the concepts a model has implicitly learned, the training distribution it reflects, and potential sources of bias. We evaluate UNBOX on ImageNet-1K, Waterbirds, and CelebA through semantic fidelity tests, visual-feature correlation analyses and slice-discovery auditing. Despite operating under the strictest black-box constraints, UNBOX performs competitively with state-of-the-art white-box interpretability methods. This demonstrates that meaningful insight into a model's internal reasoning can be recovered without any internal access, enabling more trustworthy and accountable visual recognition systems.

模型可解释性黑箱分析语言模型视觉识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。