用生物启发的视觉修复机制提升视频质量评估精度
EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment
- 双分支架构分别处理全局美感与局部结构,模拟自修复过程
- 在5个公开数据集上表现优于或媲美现有方法,准确率提升显著
- 适合关注可解释性与感知建模的视频质量研究者
自由能引导的自修复机制在图像质量评估中表现优异,但在视频质量评估(VQA)中仍鲜有探索,因视频具有更复杂的时空特性且模型受限于预训练主干网络。本文提出EyeSimVQA框架,采用双分支设计:美学分支进行全局感知评估,技术分支进行细粒度结构与语义分析。各分支针对不同输入(全帧图像与分块片段)集成专用增强模块,模拟自适应修复行为。同时提出一种不破坏主干网络的高层特征融合策略,并设计生物启发的预测头,模拟扫视眼动以融合全局与局部表征。在五个公开VQA基准测试中,EyeSimVQA性能达到或超越当前最优水平,且具备更好的可解释性。
原文摘要 · Abstract (English)
Free-energy-guided self-repair mechanisms have shown promising results in image quality assessment (IQA), but remain under-explored in video quality assessment (VQA), where temporal dynamics and model constraints pose unique challenges. Unlike static images, video content exhibits richer spatiotemporal complexity, making perceptual restoration more difficult. Moreover, VQA systems often rely on pre-trained backbones, which limits the direct integration of enhancement modules without affecting model stability. To address these issues, we propose EyeSimVQA, a novel VQA framework that incorporates free-energy-based self-repair. It adopts a dual-branch architecture, with an aesthetic branch for global perceptual evaluation and a technical branch for fine-grained structural and semantic analysis. Each branch integrates specialized enhancement modules tailored to distinct visual inputs-resized full-frame images and patch-based fragments-to simulate adaptive repair behaviors. We also explore a principled strategy for incorporating high-level visual features without disrupting the original backbone. In addition, we design a biologically inspired prediction head that models sweeping gaze dynamics to better fuse global and local representations for quality prediction. Experiments on five public VQA benchmarks demonstrate that EyeSimVQA achieves competitive or superior performance compared to state-of-the-art methods, while offering improved interpretability through its biologically grounded design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。