用交互式评估框架提升文生图模型的故障发现能力
Interactive Visual Assessment for Text-to-Image Generation Models
- 通过LLM动态生成多样化文本输入,实时探测模型能力边界
- 相比传统方法,可多发现2.56倍的生成错误,包括罕见难题
- 适合模型开发者和评测人员,用于挖掘复杂失败模式
视觉生成模型在计算机图形学应用中取得显著进展,但在实际部署中仍面临挑战。现有评估方法通常采用孤立的三阶段流程:输入收集、模型生成和用户评估,存在覆盖固定、难度不变、数据泄露风险等问题,难以全面评估日益复杂的生成模型。为此,我们提出DyEval——一个基于LLM的动态交互式视觉评估框架,支持人与生成模型协同评估文生图系统。DyEval提供直观界面,让用户交互探索并分析模型行为,同时根据反馈自适应生成分层、细粒度且多样的文本输入,持续探测模型能力边界。此外,为提供可解释分析以改进模型,我们开发了上下文反思模块,利用LLM逻辑推理能力挖掘测试输入的失败触发因素,揭示模型潜在失败模式。定性与定量实验表明,DyEval能有效帮助用户识别最多比传统方法多2.56倍的生成错误,并发现复杂罕见问题,如代词生成错误和特定文化语境生成缺陷。该框架为改进生成模型提供了宝贵洞见,对提升视觉生成系统的可靠性与能力具有广泛意义。
原文摘要 · Abstract (English)
Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an isolated three-phase framework: test input collection, model output generation, and user assessment. These fashions suffer from fixed coverage, evolving difficulty, and data leakage risks, limiting their effectiveness in comprehensively evaluating increasingly complex generation models. To address these limitations, we propose DyEval, an LLM-powered dynamic interactive visual assessment framework that facilitates collaborative evaluation between humans and generative models for text-to-image systems. DyEval features an intuitive visual interface that enables users to interactively explore and analyze model behaviors, while adaptively generating hierarchical, fine-grained, and diverse textual inputs to continuously probe the capability boundaries of the models based on their feedback. Additionally, to provide interpretable analysis for users to further improve tested models, we develop a contextual reflection module that mines failure triggers of test inputs and reflects model potential failure patterns supporting in-depth analysis using the logical reasoning ability of LLM. Qualitative and quantitative experiments demonstrate that DyEval can effectively help users identify max up to 2.56 times generation failures than conventional methods, and uncover complex and rare failure patterns, such as issues with pronoun generation and specific cultural context generation. Our framework provides valuable insights for improving generative models and has broad implications for advancing the reliability and capabilities of visual generation systems across various domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。