arXiv:2601.06169cs.CV2026-01

无需训练,通过提示引导与查询对比解码提升文本生成图像的准确性

Think Bright, Diffuse Nice: Enhancing T2I-ICL via Inductive-Bias Hint Instruction and Query Contrastive Decoding

  • 用轻量提示注入任务先验知识,防止模型偏离上下文规则
  • 通过对比完整输入与省略查询的解码分布,减少幻觉生成
  • 兼容多种模型和提示设计,适合追求高效可靠的图像生成场景

文本到图像的上下文学习(T2I-ICL)通过交错的图文示例实现定制化图像生成,但面临合规失败与先验主导幻觉两大相互强化的瓶颈,形成恶性循环导致生成质量下降。现有方法依赖特定训练,限制灵活性且增加部署成本。为此,我们提出TBDN——一种无需训练的框架,集成两种互补的闭环机制:提示引导(HI)与查询对比解码(QCD)。HI通过轻量级提示工程注入任务感知的归纳偏置,使模型锚定在上下文映射规则上,缓解合规失败。QCD通过对比全输入与查询缺失的解码分布,调整语言模型的解码策略,抑制先验主导的幻觉。TBDN在CoBSAT和Text-to-Image Fast Mini-ImageNet上达到当前最优性能,对不同模型主干、提示设计和超参数具有强泛化能力。同时在Dreambench++上保持出色的概念保留与提示遵循表现。通过打破两大瓶颈,TBDN建立了一个简洁而有效的高效可靠T2I-ICL框架。

原文摘要 · Abstract (English)

Text-to-Image In-Context Learning (T2I-ICL) enables customized image synthesis via interleaved text-image examples but faces two mutually reinforcing bottlenecks, compliance failure and prior-dominated hallucination, that form a vicious cycle degrading generation quality. Existing methods rely on tailored training, which limits flexibility and raises deployment costs. To address these challenges effectively, we propose TBDN, a training-free framework integrating two complementary closed-loop mechanisms: Hint Instruction (HI) and Query Contrastive Decoding (QCD). HI injects task-aware inductive bias via lightweight prompt engineering to anchor models on contextual mapping rules, thereby mitigating compliance failure. QCD adjusts the decoding distributions of language models by contrasting full-input and query-omitted distributions, suppressing prior-dominated hallucination. TBDN achieves State-of-the-Art performance on CoBSAT and Text-to-Image Fast Mini-ImageNet, with robust generalization across model backbones, prompt designs, and hyperparameters. It also maintains promising performance in concept preservation and prompt following on Dreambench++. By breaking the two bottlenecks, TBDN establishes a simple yet effective framework for efficient and reliable T2I-ICL.

文本生成图像上下文学习提示工程去幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。