用贝叶斯思想让机器人从猫的行为推断意图,提升准确率并减少误判。
Context as Prior: Bayesian-Inspired Intent Inference for Non-Speaking Agents with a Household Cat Testbed

- 将环境信息当作先验约束,结合动作和声音证据推断意图
- 在猫咪数据集上达到77.72%准确率,优于传统融合方法
- 有效抑制了因环境线索导致的错误预测,适合非语言智能体研究
现实世界中许多实体(如宠物猫、婴幼儿)无法通过语言表达目标,其意图需从丰富上下文中的不完整行为中推断。此过程存在核心模糊性:行为观测常含噪声或信息不足,而上下文虽提供强先验但易引发机械性误判。本文提出CatSignal,一种受贝叶斯启发的多模态意图推断框架,将空间上下文建模为类先验约束,行为动态与声学线索作为证据。采用上下文门控的Expert乘积结构计算后验式意图分布。在家庭猫咪多模态数据集上,留一视频测试下,该方法实现77.72%的整体准确率,优于特征拼接(71.83%)与更强的晚期融合基线。更重要的是,在模糊场景中显著降低由上下文驱动的捷径预测失败率。尽管简单融合策略在宏观F1与选择性预测上仍具竞争力,本模型在整体准确率与上下文捷径抑制方面表现最优。
原文摘要 · Abstract (English)
Many agents in real-world environments cannot reliably communicate their goals through language, including household pets, pre-verbal infants, and other non-speaking embodied agents. In such settings, intent must be inferred from incomplete behavioral observations in context-rich environments. This creates a core ambiguity: observable behavior is often noisy or underspecified, while context provides strong prior information but can also induce brittle shortcut predictions if used naively. We present CatSignal, a Bayesian-inspired probabilistic framework for multimodal intent inference that models spatial context as a prior-like constraint and behavioral observations as evidence. Rather than treating context as an ordinary input feature, our method uses a context-gated Product-of-Experts formulation to compute posterior-like intent distributions from context, pose dynamics, and acoustic cues. We instantiate this formulation in a household cat setting as a focused proof-of-concept for intent inference in non-speaking agents. Under Leave-One-Video-Out evaluation on a multimodal domestic cat dataset, the proposed prior-guided fusion achieves the best overall accuracy of 77.72%, outperforming feature concatenation (71.83%) and stronger late-fusion baselines. More importantly, it substantially reduces context-driven shortcut failures in ambiguous cases. While simpler fusion strategies remain competitive in Macro-F1 and selective prediction, the proposed model provides the strongest overall accuracy and the best suppression of context-based shortcut collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。