arXiv:2603.04419cs.CLcs.AI2026-03

视觉语言模型的物体可用性会随上下文大幅变化,影响机器人决策。

Context-Dependent Affordance Computation in Vision-Language Models

  • 分析了不同场景下模型对物体用途的判断差异,发现90%描述依赖上下文。
  • 跨模型验证显示效果一致,语义层面仍有58.5%的变化,说明深层理解也受影响。
  • 揭示了厨房与儿童移动等特定情境下的潜在语义结构,适合机器人研发者参考。

我们研究了视觉语言模型中上下文依赖的物体可用性计算现象。以Qwen3-VL-30B-A3B为基础,使用COCO-2017中3,213组场景-上下文对(479张图像,7种代理人格)进行主实验,并在LLaVA-1.5-13B上复现。结果显示显著的可用性漂移:不同上下文间平均杰卡德相似度为0.095(95%置信区间[0.092, 0.097],基于479张图像,9,244个对比对,p < 0.0001),表明超过90%的词汇描述依赖上下文;LLaVA复现结果同样显著(均值J = 0.160,84%依赖上下文)。句级余弦相似度证实语义层面存在漂移(均值=0.415,58.5%依赖上下文)。通过2,384次随机基线推理(4种温度,5个种子)排除生成噪声影响,证明这是真实上下文效应。利用托克尔分解与自助稳定性分析(1,000次重采样),识别出稳定的正交潜在因子:一个仅限厨师情境的“烹饪流形”和横跨儿童移动性的“访问轴”。词汇与语义层面差异表明,表面用词变化大于深层意义。该发现提示机器人应采用动态、查询相关的本体投影(即时本体),而非静态世界建模。本文不主张处理顺序或架构优先性,因这需超越输出行为的内部表征分析。

原文摘要 · Abstract (English)

We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-VL-30B-A3B ($n = 3{,}213$ scene-context pairs from COCO-2017: 479 images under 7 agentic personas), with a cross-model replication on LLaVA-1.5-13B. We demonstrate substantial affordance drift: mean Jaccard similarity between context conditions is $0.095$ (95% CI $[0.092, 0.097]$ across $N = 479$ images; $9{,}244$ prime pairs; $p < 0.0001$), indicating that more than 90% of lexical scene description is context-dependent; the LLaVA replication reproduces the effect (mean $J = 0.160$, 84% context-dependent). Sentence-level cosine similarity confirms drift at the semantic level (mean $= 0.415$, 58.5% context-dependent). Stochastic baseline experiments ($2{,}384$ inference runs across 4 temperatures and 5 seeds) confirm this reflects genuine context effects rather than generation noise: within-prime variance is substantially lower than cross-prime variance across all conditions. Tucker decomposition with bootstrap stability analysis ($n = 1{,}000$ resamples) reveals stable orthogonal latent factors: a "Culinary Manifold" isolated to chef contexts and an "Access Axis" spanning child-mobility contrasts. The gap between lexical (90%) and semantic (58.5%) measures indicates that surface vocabulary changes more than underlying meaning under context shifts. These findings suggest a direction for robotics: dynamic, query-dependent ontological projection (JIT Ontology) rather than static world modeling. We do not claim to establish processing order or architectural primacy; such claims require internal representational analysis beyond output behavior.

视觉语言模型上下文依赖机器人决策本体投影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。