让冻结的视觉语言模型对齐人类偏好,无需重新训练。
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
- 用三阶段后处理流程提取可解释评价维度并校准输出。
- 在六类感知任务中达72.2%准确率,优于基线11个百分点。
- 全程不修改模型权重,结果可解释,适合城市评估场景。
视觉语言模型(VLMs)能详细描述城市场景,但在安全评估和美学评价等特定任务中,难以生成可靠的偏好标签。传统方法如微调或基于人类反馈的强化学习需大量标注数据和重训练。本文提出不同思路:能否在不修改任何权重的前提下,使冻结的VLM对齐人类偏好?核心洞察是:VLM虽强于概念提取,但弱于决策校准。为此,我们设计三阶段后处理框架:(i) 从共识样本中自动挖掘可解释的评价维度;(ii) 通过观察者-辩论者-裁判链从冻结的VLM中提取稳健的概念得分;(iii) 在混合流形上使用局部加权岭回归将这些得分校准至人类评分。该方法名为UrbanAlign,在Place Pulse 2.0数据集上实现72.2%准确率(kappa=0.45),覆盖六个感知类别,较所有基线提升11.0个百分点,零样本VLM提升15.5个百分点,且保持完全可解释性与零权重修改。
原文摘要 · Abstract (English)
Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety assessment and aesthetic evaluation. The standard fix, fine-tuning or RLHF, requires large-scale annotations and model retraining. We ask a different question: can a frozen VLM be aligned with human preferences without modifying any weights? Our key insight is that VLMs are strong concept extractors but poor decision calibrators. We propose a three-stage post-hoc pipeline that exploits this asymmetry: (i) interpretable evaluation dimensions are automatically mined from consensus exemplars; (ii) an Observer-Debater-Judge chain extracts robust concept scores from the frozen VLM; and (iii) locally-weighted ridge regression on a hybrid manifold calibrates these scores to human ratings. Applied as UrbanAlign on Place Pulse 2.0, the framework reaches 72.2% accuracy (kappa=0.45) across six perception categories, outperforming all baselines by +11.0 pp and zero-shot VLM by +15.5 pp, with full interpretability and zero weight modification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。