提出人因导向的多模态复杂度模型,解析真实场景中视觉空间认知负担。
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision

- 基于具身认知理论,构建包含五类属性的多模态复杂度分类体系
- 在日常驾驶场景中验证该模型对视觉空间复杂度的表征能力
- 为自动驾驶数据集设计与人眼视觉研究提供可解释的分析框架
我们提出一种新型多模态数据分析框架,涵盖视觉、听觉与空间刺激,强调在动态自然环境中具身感知与交互中的复杂性。基于具身认知与主动视觉理论,认为具身感知复杂性源于智能体与环境的动态互动,需整体分析其定性与定量属性,如视觉空间与听觉特征。在既有视觉复杂性研究基础上,扩展出量化、结构、动态、听觉与交互五类复杂性属性,共同刻画多模态复杂性。通过实证表明该模型能有效表征日常驾驶情境下的视觉空间复杂度及其交互机制。同时讨论了该框架在创建与评估以认知人因为核心的基准数据集(如驾驶场景)中的应用潜力,以及系统研究视觉空间复杂度对人类主动视觉影响的路径。该框架为从以人为本视角自动解析3D动态环境中的复杂性奠定了基础,提供了一种以认知人因为导向的语义模板,支持可解释的计算分析。
原文摘要 · Abstract (English)
We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of complexity in embodied perception and interaction in dynamic, naturalistic settings. Grounded in theories of embodied cognition and active vision, we argue that embodied perceptual complexity emerges from an agent's dynamic engagement with the environment and must be analyzed holistically, as a combination of qualitative and quantitative attributes pertaining to, for instance, visuospatial and auditory features. Building on previous work on visual complexity, we expand this into a categorization of diverse complexity attributes -- quantitative, structural, dynamic, auditory, and interactional -- that together characterize multimodal complexity. We demonstrate how this model provides a theoretical framework for characterizing aspects of visuospatial complexity and their interactions, specifically in the context of everyday driving. We also discuss practical applications of the proposed model for creating and evaluating benchmark datasets (e.g., in driving) that centralize cognitive human factors, as well as applications aimed at systematically investigating the effect of visuospatial complexity on human active vision from the viewpoint of visual perception research. The proposed framework lays the foundation for automated methods that interpret complexity in 3D dynamic environments from a human-centered perspective, serving as a semantic template for explainable computational analysis of visuospatial complexity with a categorical focus on cognitive human factors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。