arXiv:2606.09110cs.CV2026-06

用智能代理动态选策略,减少动态场景的HDR伪影。

HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging

论文配图:HDRAgent: An Agentic Framework for Multi-Exposure HDR Imaging
图 1 · 摘自论文原文
  • 引入多模态大模型感知场景,动态匹配历史案例和工具知识
  • 通过反馈机制积累质量评估结果,优化后续策略选择
  • 基于大模型解析动态区域,实现无参考帧的生成对齐

现有大多数多曝光HDR方法采用固定前馈重建范式,在复杂动态场景中容易产生鬼影伪影。为此,我们提出HDRAgent,首个面向HDR成像的智能体驱动框架,能根据当前场景条件自适应选择重建策略。为提供场景特定先验,我们设计细粒度上下文知识匹配(FCM)模块,利用多模态大语言模型(MLLM)生成的场景感知能力,检索相关历史案例与工具知识,并组织为结构化证据,用于支持MLLM驱动的自适应工具调度。此外,我们提出感知-失真反馈机制,将执行后的质量评估与伪影诊断转化为结构化反馈,存入历史记忆,以辅助后续上下文知识的迭代优化与策略选择。针对极端运动导致对齐失效的问题,设计了代理引导的生成式对齐策略,通过MLLM进行动态区域解析,在参考帧引导下重构非参考帧中不可靠内容。实验表明,HDRAgent有效降低鬼影与局部伪影,同时在客观指标与视觉质量上达到或优于现有方法。

原文摘要 · Abstract (English)

Most existing multi-exposure HDR methods follow a fixed feed-forward reconstruction paradigm, making them prone to ghosting artifacts in complex dynamic scenes. To address this issue, we propose HDRAgent, the first agent-driven framework for HDR imaging, which adaptively selects reconstruction strategies according to the current scene conditions. Specifically, to provide scene-specific prior knowledge, we introduce a fine-grained contextual knowledge matching (FCM) module. This module leverages multimodal large language model (MLLM)-derived scene perception to retrieve relevant historical cases and tool knowledge, organizing them into structured evidence for MLLM-based adaptive tool scheduling. In addition, we propose a perception--distortion feedback mechanism that transforms post-execution quality assessment and artifact diagnosis into structured feedback, which is accumulated in historical memory to help subsequent contextual knowledge refinement and strategy selection. Furthermore, considering that extreme motion can invalidate alignment methods, we design an agent-guided generative alignment strategy that uses MLLM-based dynamic-region parsing to reconstruct unreliable contents in non-reference frames under reference-frame guidance. Experiments demonstrate that HDRAgent effectively reduces ghosting and local artifacts while achieving competitive or superior objective performance and visual quality.

HDR成像智能代理多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。