小改动就能让健康类社交媒体分析出错,新防御方法有效守住数据真实性
From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse
- 用微小干扰篡改支持样本,破坏大模型的上下文学习能力
- 攻击使情感判断错误率高达67%,防御后准确率稳定在46.7%
- 适合关注医疗舆情监控安全性的研究者与从业者参考
本研究探讨了大型语言模型在上下文学习(ICL)中,面对社交媒体公共卫生情绪分析场景时,易受数据投毒攻击的影响。以人类偏肺病毒(HMPV)相关推文为例,通过同义词替换、否定插入和随机扰动等轻微扰动,对支持样本进行攻击。即使这些微小修改,也导致高达67%的案例中情感标签发生翻转。为应对该问题,引入谱特征防御(Spectral Signature Defense),可有效剔除被污染样本,同时保持原始语义与情感一致性。防御后,ICL准确率稳定在约46.7%,逻辑回归验证准确率达100%,表明该方法成功维护了数据集完整性。研究将此前关于ICL投毒的理论发现拓展至高风险的公共卫生话语分析实际场景,揭示了上下文学习在攻击下的脆弱性,并凸显谱防御在提升健康类社交监测系统可靠性方面的价值。
原文摘要 · Abstract (English)
This study explored how in-context learning (ICL) in large language models can be disrupted by data poisoning attacks in the setting of public health sentiment analysis. Using tweets of Human Metapneumovirus (HMPV), small adversarial perturbations such as synonym replacement, negation insertion, and randomized perturbation were introduced into the support examples. Even these minor manipulations caused major disruptions, with sentiment labels flipping in up to 67% of cases. To address this, a Spectral Signature Defense was applied, which filtered out poisoned examples while keeping the data's meaning and sentiment intact. After defense, ICL accuracy remained steady at around 46.7%, and logistic regression validation reached 100% accuracy, showing that the defense successfully preserved the dataset's integrity. Overall, the findings extend prior theoretical studies of ICL poisoning to a practical, high-stakes setting in public health discourse analysis, highlighting both the risks and potential defenses for robust LLM deployment. This study also highlights the fragility of ICL under attack and the value of spectral defenses in making AI systems more reliable for health-related social media monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。