评测图像文本真伪验证系统,推动可信AI发展
The Automatic Verification of Image-Text Claims (AVerImaTeC) Shared Task
- 设计图像文本真伪验证共享任务,支持外部搜索与知识库
- 测试阶段所有系统超越基线,冠军得分0.5455
- 适合关注多模态事实核查的科研与工程人员
自动图像文本声明验证(AVerImaTeC)共享任务旨在推动真实世界图像文本声明证据检索与验证系统的发展。参赛者可使用外部知识源(如网络搜索引擎)或组织方提供的结构化知识库。系统性能通过AVerImaTeC评分衡量,该评分定义为条件判断准确率:仅当相关证据得分超过预设阈值时,判断才视为正确。开发阶段共收到14个提交,测试阶段有6个提交。测试阶段所有系统均优于基线。冠军团队HUMANE取得0.5455的AVerImaTeC分数。本文详细描述了该共享任务,呈现完整评估结果,并讨论关键洞察与经验教训。
原文摘要 · Abstract (English)
The Automatic Verification of Image-Text Claims (AVerImaTeC) shared task aims to advance system development for retrieving evidence and verifying real-world image-text claims. Participants were allowed to either employ external knowledge sources, such as web search engines, or leverage the curated knowledge store provided by the organizers. System performance was evaluated using the AVerImaTeC score, defined as a conditional verdict accuracy in which a verdict is considered correct only when the associated evidence score exceeds a predefined threshold. The shared task attracted 14 submissions during the development phase and 6 submissions during the testing phase. All participating systems in the testing phase outperformed the baseline provided. The winning team, HUMANE, achieved an AVerImaTeC score of 0.5455. This paper provides a detailed description of the shared task, presents the complete evaluation results, and discusses key insights and lessons learned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。