arXiv:2506.00930cs.AIcs.CL2025-06ACL被引 2

让视觉语言模型适应不同人的个性化认知,提升真实场景辅助能力

Aligning VLM Assistants with Personalized Situated Cognition

  • 基于角色集合概念建模个体认知差异,设计个性化对齐框架
  • 构建包含18000条数据、20名不同角色个体的基准测试集
  • 通过动作评估与奖励模型实现认知感知的个性化响应

以通用人类目标(如无害性、无幻觉)对齐的视觉语言模型(VLM)已成为人类处理视觉任务的有用助手。然而,不同背景的人在相同情境下认知各异,对VLM助手的期望也存在个性化差异。这凸显了将VLM助手与个性化情境认知对齐的紧迫性。为此,我们首先基于社会学中的角色集合(Role-Set)概念简化个体表征,并提出通过评估个体行为来检验个性化对齐是否达成。进一步,我们构建了一个名为PCogAlignBench的基准测试,包含18,000个实例和20位具有不同角色集合的个体。最后,我们提出一种名为PCogAlign的框架,其构建了认知感知且基于行为的奖励模型,实现个性化对齐。实验结果与人工评估验证了PCogAlignBench的可靠性及所提方法的有效性。相关代码与基准数据集将在https://github.com/NLPGM/PCogAlign开源。

原文摘要 · Abstract (English)

Vision-language models (VLMs) aligned with general human objectives, such as being harmless and hallucination-free, have become valuable assistants of humans in managing visual tasks. However, people with diversified backgrounds have different cognition even in the same situation. Consequently, they may have personalized expectations for VLM assistants. This highlights the urgent need to align VLM assistants with personalized situated cognition for real-world assistance. To study this problem, we first simplify it by characterizing individuals based on the sociological concept of Role-Set. Then, we propose to evaluate the individuals' actions to examine whether the personalized alignment is achieved. Further, we construct a benchmark named PCogAlignBench, which includes 18k instances and 20 individuals with different Role-Sets. Finally, we present a framework called PCogAlign, which constructs a cognition-aware and action-based reward model for personalized alignment. Experimental results and human evaluations demonstrate the reliability of the PCogAlignBench and the effectiveness of our proposed PCogAlign. We will open-source the constructed benchmark and code at https://github.com/NLPGM/PCogAlign.

视觉语言模型个性化对齐认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。