通过多智能体推理提升图像事件背景理解能力
GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
- 采用分层多智能体架构,融合时空地理信息进行深度推理
- 在新指标GREAT下,显著提升图像上下文理解准确率
- 适合媒体、教育领域需深度分析图像背景的场景
来自重大事件的公开图像蕴含重要上下文信息,对新闻报道与教育具有关键价值。然而现有方法难以准确提取此类相关信息。为此,我们提出GETReason(地理事件时间推理)框架,突破表面描述,深入推断图像的全局事件、时间和地理背景信息,以增强理解。同时,我们引入GREAT(地理推理与事件时序对齐准确性)评估指标,用于衡量基于推理的图像理解性能。通过分层多智能体方法,结合加权推理评估,实验表明可有效挖掘图像背后的深层意义,实现图像与事件背景的精准关联。
原文摘要 · Abstract (English)
Publicly significant images from events hold valuable contextual information, crucial for journalism and education. However, existing methods often struggle to extract this relevance accurately. To address this, we introduce GETReason (Geospatial Event Temporal Reasoning), a framework that moves beyond surface-level image descriptions to infer deeper contextual meaning. We propose that extracting global event, temporal, and geospatial information enhances understanding of an image's significance. Additionally, we introduce GREAT (Geospatial Reasoning and Event Accuracy with Temporal Alignment), a new metric for evaluating reasoning-based image understanding. Our layered multi-agent approach, assessed using a reasoning-weighted metric, demonstrates that meaningful insights can be inferred, effectively linking images to their broader event context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。