用大模型提升小视觉模型的人类知识对齐能力
LVLM-Aided Alignment of Task-Specific Vision Models
- 利用大视觉语言模型双向翻译模型行为与人类指令
- 显著降低模型对虚假特征和群体偏见的依赖
- 无需细粒度反馈,适合领域专家快速调试
在高风险领域,小型任务特定视觉模型因计算成本低且可解释性强而至关重要。然而,这些模型的解释常显示其与人类领域知识不一致,反而依赖虚假相关性,导致实际部署时表现脆弱。为此,我们提出一种新方法——基于大视觉语言模型(LVLM)的视觉对齐(LVLM-VA),通过双向接口将模型行为转化为自然语言,并将人类层级规范映射为图像级批评,实现领域专家与模型的有效交互。在合成与真实数据集上的验证表明,该方法显著提升了模型行为与人类规范的一致性,有效减少了对虚假特征和群体特定偏见的依赖,且无需细粒度反馈。
原文摘要 · Abstract (English)
In high-stakes domains, small task-specific vision models are crucial due to their low computational requirements and the availability of numerous methods to explain their results. However, these explanations often reveal that the models do not align well with human domain knowledge, relying instead on spurious correlations. This might result in brittle behavior once deployed in the real-world. To address this issue, we introduce a novel and efficient method for aligning small task-specific vision models with human domain knowledge by leveraging the generalization capabilities of a Large Vision Language Model (LVLM). Our LVLM-Aided Visual Alignment (LVLM-VA) method provides a bidirectional interface that translates model behavior into natural language and maps human class-level specifications to image-level critiques, enabling effective interaction between domain experts and the model. Our method demonstrates substantial improvement in aligning model behavior with human specifications, as validated on both synthetic and real-world datasets. We show that it effectively reduces the model's dependence on spurious features and on group-specific biases, without requiring fine-grained feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。