arXiv:2607.22352cs.CVcs.AI2026-07

从残留痕迹反推刚发生的人-环境互动,开启新场景理解范式。

Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions

论文配图:Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions
图 1 · 摘自论文原文
  • 基于热、紫外、可见光多模态痕迹逆向推断事件
  • 在三分钟内可重建坐、触、移动等动作的合理历史画面
  • 适合对物理推理与生成模型融合感兴趣的学者

我们提出时间反转成像,一种从衰减的多模态痕迹中推断场景近期事件的新范式。不同于视频插值或外推,目标是通过热、紫外和可见光谱中可观测的残余物理印记,还原人-环境交互的历史。为此,我们构建了首个概念验证数据集TRACE-HEI,包含同步的三模态视频序列,记录坐、触碰、移动物体、液体泼洒等动作,覆盖多种材质,最长记录间隔达三分钟。为建立基准,我们提出一种多模态推理方法:提取检测到痕迹的结构化文本描述,并用于约束视觉-语言引导的扩散模型,以重建合理的过去帧。实验表明,在互补模态减少解歧义的前提下,从衰减痕迹中推断近期事件虽具挑战性但可行。该工作奠定了时间反转成像的首个计算与实验基础,连接视觉、物理与生成推理,拓展了超越瞬时观测的场景理解方向。

原文摘要 · Abstract (English)

We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or interpolating video frames, our goal is to infer past human-environment interactions from residual physical imprints observable in thermal, ultraviolet, and visible spectra. To study this problem, we present TRACE-HEI, the first proof-of-concept dataset for time-reversed imaging, containing synchronized tri-modal video sequences of actions such as sitting, touching, moving objects, and liquid spills, captured across diverse materials and recorded up to three minutes after contact. To establish the benchmark, we propose a multimodal inference approach that extracts structured textual descriptions of detected traces and uses them to constrain a vision-language-guided diffusion model for reconstructing plausible past frames. Experiments show that inferring recent events from fading traces is challenging but feasible when complementary modalities reduce solution ambiguity. This work defines the first computational and experimental foundation for time-reversed imaging, bridging vision, physics, and generative reasoning, and opening new directions for scene understanding beyond instantaneous observation.

时间反转多模态感知生成建模场景重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。