用AI预测城市级食品安全风险,数据少也能准
Leveraging AI for fine-grained food safety risk forecasting in sparse data conditions

- 用Transformer融合1100万条检查数据与社会经济环境信息
- 在2022年数据上显著优于基线模型,提升检测率
- 适合政府监管、公共卫生决策者参考
保障食品安全是重大公共卫生挑战,尤其在检查资源有限、区域采样数据稀疏的情况下。本研究提出一种基于Transformer的框架,通过整合超过1100万条检查记录,并融合来自统计年鉴的人口、经济与环境指标,实现精细化的城市级食品安全风险预测。采用三阶段预训练设计,利用威尔逊区间(Wilson interval)中的部分监督信号(涵盖安全与风险排序),结合半监督标签优化,有效利用历史数据,即使本地样本量不足也能建模。2022年数据实验表明,该方法显著优于基线模型。后续与浙江省市场监督管理局合作的实地实验显示,相较人工规划方案,该方法提升了检测率并优化了检查资源配置。监管决策观察发现,检查员采用阈值启发式策略,提示可进一步通过训练或决策支持界面增强AI风险评分的实际影响。总体而言,大规模公共检查数据、基于威尔逊区间的置信建模与先进深度学习的有机结合,有助于更早、更精细地识别食品安全威胁,减少对被动应对的依赖,推动全球食品供应链的主动、数据驱动监管。
原文摘要 · Abstract (English)
Ensuring food safety represents a critical public health challenge, particularly when inspection resources are limited and regional sampling data are sparse. This study proposes a Transformer-based framework capable of forecasting fine-grained, city-level food safety risks by unifying over 11 million inspection records with supplemental demographic, economic, and environmental indicators extracted from the Statistical Yearbook. A three-stage pretraining design leverages partial supervision from the Wilson interval (capturing both safety and risk rankings), together with semi-supervised label refinement, to effectively utilize historical records even when local sample sizes are insufficient. Experimental evaluations on data from 2022 show that the proposed approach outperforms baselines significantly. A subsequent field experiment in collaboration with the Zhejiang Provincial Administration for Market Regulation further demonstrates improved detection rates and more efficient allocation of inspection resources compared to a manually developed plan. Observations of regulatory decision-making reveal a threshold-based heuristic employed by inspectors, hinting that additional training or decision-support interfaces could further enhance the impact of AI-generated risk scores. Overall, these findings underscore that a rigorous integration of large-scale public inspection data, Wilson interval-based confidence modeling, and advanced deep learning can facilitate earlier and more granular identification of food safety threats. By reducing reliance on reactive measures alone, the proposed framework has the potential to advance proactive, data-driven oversight of the global food supply.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。