模拟人类视听搜索行为,实现高效资源分配。
Simulating Human Audiovisual Search Behavior
- 基于资源理性决策构建具身视听搜索模型
- 准确复现人类在复杂场景下的搜索耗时与努力变化
- 适用于人机交互界面优化设计
在不确定环境下,结合听觉与视觉线索定位目标(如在拥挤停车场找车或虚拟会议中识别发言者)需要权衡努力、时间与准确性。现有视听搜索模型通常将感知与行动分开处理,忽略了人类如何自适应协调身体移动与感官策略。我们提出Sensonaut,一个具身视听搜索的计算模型。核心假设是:人们会以最有效方式部署身体与感官系统,以提升定位目标的概率,同时在感知限制下权衡时间和努力。模型将此建模为部分可观测环境下的资源理性决策问题。通过新收集的人类数据验证,模型能复现人类在任务复杂度、遮挡和干扰条件下的搜索时间与努力的自适应调整,以及典型的人类错误模式。该模拟结果有助于设计降低搜索成本与认知负荷的视听交互界面。
原文摘要 · Abstract (English)
Locating a target based on auditory and visual cues$\unicode{x2013}$such as finding a car in a crowded parking lot or identifying a speaker in a virtual meeting$\unicode{x2013}$requires balancing effort, time, and accuracy under uncertainty. Existing models of audiovisual search often treat perception and action in isolation, overlooking how people adaptively coordinate movement and sensory strategies. We present Sensonaut, a computational model of embodied audiovisual search. The core assumption is that people deploy their body and sensory systems in ways they believe will most efficiently improve their chances of locating a target, trading off time and effort under perceptual constraints. Our model formulates this as a resource-rational decision-making problem under partial observability. We validate the model against newly collected human data, showing that it reproduces both adaptive scaling of search time and effort under task complexity, occlusion, and distraction, and characteristic human errors. Our simulation of human-like resource-rational search informs the design of audiovisual interfaces that minimize search cost and cognitive load.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。