arXiv:2601.16811cs.CVcs.AI2026-01被引 2

用眼动数据提升模型对室内美学体验的预测能力

Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors

  • 双分支模型融合视觉与眼动信号进行美学评估
  • 主观维度预测准确率达66.8%,客观维度达72.2%
  • 眼动信息可提升模型泛化力,适合设计类应用

理解人们对室内空间的感知与评价对促进健康设计至关重要。然而,由于感知主观性强、视觉反应复杂,美学体验预测仍具挑战。本研究提出一种双分支CNN-LSTM框架,融合视觉特征与眼动信号,预测住宅内饰的美学评价。我们收集了224段室内设计视频及28名参与者同步的眼动数据,涵盖15个美学维度的评分。所提模型在客观维度(如光线)上达到72.2%准确率,在主观维度(如放松感)上达66.8%,优于现有视频基线模型,且在主观任务中表现更优。值得注意的是,加入眼动训练的模型在仅输入视觉信息时仍保持良好性能。消融实验表明,瞳孔反应对客观评估贡献最大,而注视点与视觉线索结合则显著提升主观评价。结果表明,眼动信号作为训练中的特权信息,可增强美学评估工具的实际应用价值。

原文摘要 · Abstract (English)

Understanding how people perceive and evaluate interior spaces is essential for designing environments that promote well-being. However, predicting aesthetic experiences remains difficult due to the subjective nature of perception and the complexity of visual responses. This study introduces a dual-branch CNN-LSTM framework that fuses visual features with eye-tracking signals to predict aesthetic evaluations of residential interiors. We collected a dataset of 224 interior design videos paired with synchronized gaze data from 28 participants who rated 15 aesthetic dimensions. The proposed model attains 72.2% accuracy on objective dimensions (e.g., light) and 66.8% on subjective dimensions (e.g., relaxation), outperforming state-of-the-art video baselines and showing clear gains on subjective evaluation tasks. Notably, models trained with eye-tracking retain comparable performance when deployed with visual input alone. Ablation experiments further reveal that pupil responses contribute most to objective assessments, while the combination of gaze and visual cues enhances subjective evaluations. These findings highlight the value of incorporating eye-tracking as privileged information during training, enabling more practical tools for aesthetic assessment in interior design.

美学评估眼动追踪多模态学习室内设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。