arXiv:2605.27195cs.CL2026-05

新基准和度量方法提升疫情曲线数据提取的准确性与合理性

EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization

论文配图:EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization
图 1 · 摘自论文原文
  • 构建1000张真实疫情曲线图像库,匹配时间序列特性设计新评估指标
  • 顶尖模型仅达52.3%新指标分数,传统方法差距被显著拉大
  • 适合公共卫生、医疗数据分析及图表自动化提取研究者使用

视觉语言模型(VLMs)在图表转数据任务中的评估面临两大问题:现有基准测试头部空间已趋饱和(前沿模型在ChartQA上得分超89%),且评估指标将提取点视为无序键值对,忽略时间序列结构,对微小对齐偏差过度惩罚。为此,本文提出EpiCurveBench——一个包含1000张来自多元公共卫生来源的真实疫情曲线图像的基准集,以及EpiCurveSimilarity(ECS)评估指标,该指标通过动态规划对齐预测与真实序列,容忍局部时间偏移与缺失,并按比例惩罚。在六种方法(三种前沿闭源VLM、一个开源VLM及两个专用图表提取系统)上评估发现,最强模型仅达52.3% ECS;而传统指标(RMS、SCRM)将四类通用VLM压缩至5分区间,而ECS展现25分跨度。进一步验证显示,更高ECS值能更准确预测总数量、峰值时间与幅度误差,以及增长率保真度;其相关性较动态时间规整(DTW)高出1.5至3.6倍,因后者缺乏缺口惩罚机制,无法区分截断预测与时间对齐预测。

原文摘要 · Abstract (English)

Chart-to-data extraction with vision-language models (VLMs) is increasingly evaluated on benchmarks that show diminishing headroom (frontier VLMs exceed 89% on ChartQA) and with metrics that treat extracted points as unordered key-value pairs, ignoring the temporal structure of time series and penalizing small alignment shifts as catastrophic failures. We address both gaps with EpiCurveBench, a benchmark of 1,000 real-world epidemic curve images curated from diverse public-health sources, and EpiCurveSimilarity (ECS), an evaluation metric that aligns predicted and ground-truth series via dynamic programming, tolerating local temporal shifts and gaps while penalizing them proportionally. Evaluating six methods--three frontier closed VLMs, one open VLM, and two specialized chart-extraction systems--we find the strongest model reaches only 52.3% ECS, and that ECS spreads the four general-purpose VLMs over a 25-point range where key-value metrics (RMS, SCRM) compress them into a 5-point band. We further validate ECS against four downstream epidemiological summary statistics, finding that higher ECS predicts smaller errors in total counts, peak timing, and peak magnitude, and higher growth-rate fidelity; across all four, ECS correlates 1.5--3.6 times more strongly than Dynamic Time Warping, which lacks a gap penalty and therefore cannot distinguish a truncated prediction from a temporally faithful one. EpiCurveBench targets a high-impact public-health application--unlocking decades of outbreak data trapped in published figures--but the benchmark and metric apply directly to any structured time-series chart-extraction setting.

图表提取疫情分析多模态评估时间序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。