arXiv:2506.01308cs.CLcs.IR2025-06

用轻量模型快速识别网络公共健康担忧,助力政策响应。

A Platform for Investigating Public Health Content with Efficient Concern Classification

  • 用大模型知识迁移训练轻量分类器,高效识别文本中的健康关切
  • 分析18.6万条数据,发现重大事件前后议题频率变化趋势
  • 支持直接上传、自动抓取链接,适合公共卫生官员日常使用

近期在线内容中关于公共卫生举措的担忧增多,已阻碍全球预防措施的推广。未来公共健康工作需理解这些内容所引发的关切及其回应方式。为此,我们提出ConcernScope平台,利用教师-学生框架在大型语言模型与轻量分类器间实现知识迁移,快速有效识别文本语料中的健康关切。平台支持直接上传大规模文件、自动抓取特定网址及文本编辑。其基于公共健康关切分类体系构建,面向公共卫生官员。我们展示了多个应用场景:在社区数据集中引导式探索常见关切实例;对186,000个样本进行时间序列分析,识别关切趋势;比较重大事件前后的主题频率变化。

原文摘要 · Abstract (English)

A recent rise in online content expressing concerns with public health initiatives has contributed to already stalled uptake of preemptive measures globally. Future public health efforts must attempt to understand such content, what concerns it may raise among readers, and how to effectively respond to it. To this end, we present ConcernScope, a platform that uses a teacher-student framework for knowledge transfer between large language models and light-weight classifiers to quickly and effectively identify the health concerns raised in a text corpus. The platform allows uploading massive files directly, automatically scraping specific URLs, and direct text editing. ConcernScope is built on top of a taxonomy of public health concerns. Intended for public health officials, we demonstrate several applications of this platform: guided data exploration to find useful examples of common concerns found in online community datasets, identification of trends in concerns through an example time series analysis of 186,000 samples, and finding trends in topic frequency before and after significant events.

公共健康文本分类舆情分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。