arXiv:2602.23746cs.HCcs.CV2026-02中稿 · CHI 2026 Poster tr…

对比人类与AI在模糊日文字符识别中的决策差异。

Shape vs. Context: Examining Human--AI Gaps in Ambiguous Japanese Character Recognition

  • 用β-VAE生成连续变化的日文字形,测试单一字符识别
  • 人类与AI的判断边界不同,上下文可部分提升对齐度
  • 揭示人机认知差异,为对齐评测提供基础

高文本识别性能并不意味着视觉语言模型(VLM)在解决歧义时具备类人决策模式。我们通过直接比较人类与VLM,在使用β-VAE生成的连续插值日文字形上开展研究。在仅依赖字形的单字符识别任务中估计决策边界,并评估在词级上下文中将模糊字符置于人类决策边界附近时,VLM响应是否与人类判断对齐。结果表明,人类与VLM在形状仅任务中的决策边界存在差异;而在上下文情境下,某些条件下可改善对齐程度。这些发现揭示了质性行为差异,为构建人机对齐基准提供了基础洞见。

原文摘要 · Abstract (English)

High text recognition performance does not guarantee that Vision-Language Models (VLMs) share human-like decision patterns when resolving ambiguity. We investigate this behavioral gap by directly comparing humans and VLMs using continuously interpolated Japanese character shapes generated via a $β$-VAE. We estimate decision boundaries in a single-character recognition (shape-only task) and evaluate whether VLM responses align with human judgments under shape in context (i.e., embedding an ambiguous character near the human decision boundary in word-level context). We find that human and VLM decision boundaries differ in the shape-only task, and that shape in context can improve human alignment in some conditions. These results highlight qualitative behavioral differences, offering foundational insights toward human--VLM alignment benchmarking.

人机对齐视觉语言模型字符识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。