arXiv:2510.05408cs.CVcs.AI2025-10被引 1

用热成像反推过去画面,让时间倒流成为可能

See the past: Time-Reversed Scene Reconstruction from Thermal Traces Using Visual Language Models

  • 结合视觉语言模型与约束扩散过程,从热痕迹重建过去场景
  • 可还原长达120秒前的场景,生成符合语义的合理图像
  • 适合刑侦、监控等需要追溯事件的场景应用

从当前观测中恢复过去是极具挑战性的任务,具有法医和场景分析的应用潜力。热成像通过红外波段获取肉眼不可见的信息;由于人体温度通常高于环境(37℃-98.6°F),坐、触碰或倚靠等行为会留下残余热痕迹。这些逐渐消散的热印迹可作为被动的时间编码,用于推断远超RGB相机能力的近期事件。本文提出一种时间逆向重建框架,利用配对的RGB与热成像数据,恢复数秒前的场景状态。该方法将视觉语言模型(VLMs)与约束扩散过程相结合:一个VLM生成场景描述,另一个引导图像重建,确保语义与结构一致。在三个受控场景中评估,证明可成功重构最多120秒前的合理画面,为热痕迹驱动的时间逆向成像提供了首个实现路径。

原文摘要 · Abstract (English)

Recovering the past from present observations is an intriguing challenge with potential applications in forensics and scene analysis. Thermal imaging, operating in the infrared range, provides access to otherwise invisible information. Since humans are typically warmer (37 C -98.6 F) than their surroundings, interactions such as sitting, touching, or leaning leave residual heat traces. These fading imprints serve as passive temporal codes, allowing for the inference of recent events that exceed the capabilities of RGB cameras. This work proposes a time-reversed reconstruction framework that uses paired RGB and thermal images to recover scene states from a few seconds earlier. The proposed approach couples Visual-Language Models (VLMs) with a constrained diffusion process, where one VLM generates scene descriptions and another guides image reconstruction, ensuring semantic and structural consistency. The method is evaluated in three controlled scenarios, demonstrating the feasibility of reconstructing plausible past frames up to 120 seconds earlier, providing a first step toward time-reversed imaging from thermal traces.

热成像时间逆向视觉语言模型场景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。