将文本视为不完整约束,提升视觉语言模型在模糊描述下的鲁棒性。
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

- 用核心-残差对齐框架,从多视角描述中提取共识语义核心
- 在ImageNet上达到81.42%零样本准确率,对抗攻击下仍保持64.05%鲁棒性
- 适合需要可靠图文匹配的开放词汇识别与视觉大模型迁移任务
视觉语言对齐支持开放词汇识别、检索与大模型定位,但自然描述常信息不足,导致相似度脆弱且过度自信。本文提出文本作为部分约束(TPC)框架,将多视角描述视为不完整监督。该方法提炼共识语义核心作为对齐目标,学习单视图核心预测器用于标准推理,并显式抑制视觉-语言相似度对未言明残差的依赖。不确定性感知对比损失在描述分歧时软化对齐,减少弱语言约束下的过激更新。在零样本识别与对抗鲁棒性测试中,TPC在ImageNet上实现81.42%(清洁)/64.05%(鲁棒)Top-1准确率,在Avg-14迁移套件上达76.19%/52.03%;在LLaVA-1.5-7B架构下,提升视觉大模型迁移性能至85.16 POPE F1与59.57 OKVQA准确率。结果表明,将文本建模为部分约束是实现更可靠视觉语言表示的有效路径。
原文摘要 · Abstract (English)
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stable across caption views and whose confidence reflects how strongly text constrains an image. We propose Text as Partial Constraint (TPC), a core-residual alignment framework that treats multi-view captions as incomplete supervision. It distills a consensus semantic core as the alignment target, learns a single-view core predictor for standard inference with one query, and explicitly discourages vision-language similarity from depending on the orthogonal unsaid residual. An uncertainty-aware contrastive objective further softens alignment when caption views disagree, reducing overconfident updates under weak language constraints. Across zero-shot recognition and adversarial robustness, TPC achieves 81.42/64.05 Top-1 clean/robust accuracy on ImageNet and 76.19/52.03 on an Avg-14 transfer suite, while improving LVLM transfer with 85.16 POPE F1 and 59.57 OKVQA accuracy under an LLaVA-1.5-7B stack. These results suggest that modeling text as a partial constraint is a practical and principled route to more reliable vision-language representations under underspecified language supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。