arXiv:2511.10094cs.LGcs.CV2025-11被引 4

自动识别生成模型的物理合理性错误模式并解释原因。

How does My Model Fail? Automatic Identification and Interpretation of Physical Plausibility Failure Modes with Matryoshka Transcoders

  • 用多层编码器架构学习不同粒度的物理特征。
  • 在8个主流模型上发现多种物理错误模式,准确率优于现有方法。
  • 适合关注生成模型可靠性与可解释性的研究者使用。

尽管近期生成模型在遵循指令和生成真实输出方面表现出色,但仍易出现显著的物理合理性错误。这些错误在实际应用中至关重要,却常被现有评估方法忽略。目前尚无框架能自动识别并解释自然语言中的具体物理错误模式,阻碍了针对性改进。本文提出马特约什卡转换器(Matryoshka Transcoders),一种用于自动发现和解释生成模型中物理合理性特征的新框架。该方法将马特约什卡表示学习范式扩展至转换器架构,实现多层次稀疏特征学习。通过在物理合理性分类器的中间表示上训练,并利用大规模多模态模型进行解释,该方法无需人工特征工程即可识别多样化的物理相关错误模式,在特征相关性和准确性上均优于现有方法。我们基于发现的视觉模式建立了生成模型物理合理性评估基准。对8个先进生成模型的分析揭示了其违反物理约束的具体方式,为后续模型改进提供了重要依据。

原文摘要 · Abstract (English)

Although recent generative models are remarkably capable of producing instruction-following and realistic outputs, they remain prone to notable physical plausibility failures. Though critical in applications, these physical plausibility errors often escape detection by existing evaluation methods. Furthermore, no framework exists for automatically identifying and interpreting specific physical error patterns in natural language, preventing targeted model improvements. We introduce Matryoshka Transcoders, a novel framework for the automatic discovery and interpretation of physical plausibility features in generative models. Our approach extends the Matryoshka representation learning paradigm to transcoder architectures, enabling hierarchical sparse feature learning at multiple granularity levels. By training on intermediate representations from a physical plausibility classifier and leveraging large multimodal models for interpretation, our method identifies diverse physics-related failure modes without manual feature engineering, achieving superior feature relevance and feature accuracy compared to existing approaches. We utilize the discovered visual patterns to establish a benchmark for evaluating physical plausibility in generative models. Our analysis of eight state-of-the-art generative models provides valuable insights into how these models fail to follow physical constraints, paving the way for further model improvements.

生成模型物理合理性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。