LLM难辨印度牛尿治便秘这类文化嵌入式健康谣言
When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation

- 用多语言视频文本分析发现,谣言融合宗教话语与伪科学
- 三种主流LLM在识别文化隐含谣言时准确率普遍偏低
- 适合关注AI伦理、跨文化内容审核的研究者阅读
社交媒体已成为全球南方地区健康信息的主要渠道。以印度YouTube平台上的‘gomutra’(牛尿)言论为案例,我们通过后验的大型语言模型(LLM)辅助话语分析,研究了30段多语言视频转录文本。结果显示,推广内容将神圣传统语言与伪科学主张融合,其修辞风格甚至被去伪内容所模仿,形成难以识别的语义模式;而训练数据主要来自西方语料的LLM对此类文化嵌入式谣言系统性失能。在GPT-4o、Gemini 2.5 Pro、DeepSeek-V3.1三款模型上,不同提示语气导致分析结果差异显著,且性别化修辞与提示设计进一步加剧误判。研究指出,仅靠提示工程无法弥补LLM在文化敏感性上的根本缺陷。
原文摘要 · Abstract (English)
Social media platforms have become primary channels for health information in the Global South. Using gomutra (cow urine) discourse on YouTube in India as a case study, we present a post-facto Large Language Model (LLM)-assisted discourse analysis of 30 multilingual transcripts showing that promotional content blends sacred traditional language with pseudo-scientific claims in ways that sophisticated debunking content itself mirrors, creating a rhetorical register that LLMs, trained predominantly on Western corpora, are systematically ill-equipped to analyse. Varying prompt tone across three LLMs (GPT-4o, Gemini 2.5 Pro, DeepSeek-V3.1), we find that culturally embedded health misinformation does not look like ordinary misinformation, and this cultural obfuscation extends to gendered rhetoric and prompt design, compounding analytical unreliability. Our findings argue that cultural competency in LLM-assisted discourse analysis cannot be retrofitted through prompt engineering alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。