arXiv:2503.19800cs.CL2025-03ACL被引 26

用合成数据提升长尾食品风险文本分类效果

SemEval-2025 Task 9: The Food Hazard Detection Challenge

  • 用大模型生成合成数据缓解食品风险类别数据稀疏问题
  • 三类模型在两个子任务上均达到相近最优性能
  • 适合关注食品安全预警与长尾分类的研究者

本挑战聚焦于长尾分布下的文本型食品风险预测。任务分为两个子任务:(1) 判断网页文本是否涉及十类食品风险并识别对应食品类别;(2) 进行更细粒度分类,同时标注具体风险与产品。研究发现,大语言模型生成的合成数据对长尾类别有显著增强效果。此外,微调后的编码器-编码器、编码器-解码器及解码器-only 模型在两个子任务中表现相当。挑战期间,我们逐步发布了6,644条人工标注的食品事故报告(采用CC BY-NC-SA 4.0许可)。

原文摘要 · Abstract (English)

In this challenge, we explored text-based food hazard prediction with long tail distributed classes. The task was divided into two subtasks: (1) predicting whether a web text implies one of ten food-hazard categories and identifying the associated food category, and (2) providing a more fine-grained classification by assigning a specific label to both the hazard and the product. Our findings highlight that large language model-generated synthetic data can be highly effective for oversampling long-tail distributions. Furthermore, we find that fine-tuned encoder-only, encoder-decoder, and decoder-only systems achieve comparable maximum performance across both subtasks. During this challenge, we gradually released (under CC BY-NC-SA 4.0) a novel set of 6,644 manually labeled food-incident reports.

食品安全长尾分类合成数据文本检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。