用智能体重构犯罪现场,让每一步都可追溯、合物理、不矛盾。
VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent

- 构建智能体协调世界模型,逐条融合证据并保持可追溯性。
- 在20个测试场景中实现90.14%证据覆盖和72.17%事实一致性。
- 适合司法取证、法律AI研究者,支持多模型部署且成本仅1.82美元/场景。
世界模型能接收文本、图片、图表等多模态输入,生成符合物理规律的动态场景,为融合多源法律证据重建犯罪现场提供了新可能。然而,直接将杂乱无章的证据输入世界模型,在司法场景中会隐式丢弃证据、忽略证词矛盾,并生成违背证据记录的运动。本文提出VeriScene,一种协调世界模型的智能体:它从法医照片和不同可信度的目击陈述中重建犯罪现场,确保每个主张均可追溯至证据,每段运动均符合物理规律。VeriScene通过审计循环迭代融合证据,形成带引用的叙事;利用探针回放验证假设动态,并注入修正约束;最终生成基于融合关键帧的重演视频。在涵盖7类物理驱动案件的25个犯罪场景基准上(含139张法医风格图像和65份含虚假信息的陈述),在20个测试场景中达到0.9014的证据覆盖率和0.7217的事实一致性(0-1尺度),相比端到端多模态大模型基线提升20.35%的事实一致性与34.88%的时间连贯性,且可在四种LLM编排后端通用,单场景成本仅为1.82美元。
原文摘要 · Abstract (English)
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance with the laws of physics, thus opening a compelling application: fusing multimodal legal evidence to re-create a crime scene and re-enact how an offence could have been committed. However, feeding the raw, unorganized evidence into a world model fails in forensic use: it silently drops evidence, glosses over contradictory testimony, and produces motion that violates the evidentiary record. This paper presents VeriScene, an agent that orchestrates the world model: it reconstructs crime scenes from forensic photographs and witness statements of varying reliability, keeping every claim traceable to evidence and every motion physically plausible. VeriScene iteratively fuses the evidence into a cited narrative under an auditing loop, verifies the hypothesized dynamics via probe rollouts in the world model with corrective constraint injection, and renders the offence as a re-enactment video from a fused keyframe. On a benchmark of 25 crime scenarios across 7 physically-driven case types (139 forensic-style photographs and 65 statements with planted unreliability), VeriScene attains 0.9014 evidence coverage and 0.7217 factual consistency (0-1 scale) on the 20 test scenes, outperforming an end-to-end multimodal-LLM baseline by 20.35% in factual consistency and 34.88% in temporal coherence, while generalizing across four LLM orchestration backends at USD 1.82 per scene.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。