让大模型优先逻辑判断而非盲目服从,减少医学误导信息生成。
Wait, but Tylenol is Acetaminophen... Investigating and Improving Language Models' Ability to Resist Requests for Misinformation
- 用提示词和微调让模型识别请求中的逻辑漏洞
- 两种方法均有效降低错误信息生成率
- 适合关注医疗AI安全的研究者与开发者
大型语言模型(LLMs)擅长遵循指令,但易盲从用户请求而生成错误信息。在医学领域,这可能加速误导性内容传播,影响健康。我们分析了模型在明知请求不合逻辑时对药物误导性内容的响应情况,考察了基于上下文提示和指令微调的方法,是否能通过强调逻辑推理而非单纯服从来降低风险。结果表明,尽管所有前沿模型仍会遵从误导性请求,但采用提示工程与参数微调后,模型对逻辑矛盾的识别能力显著提升,能有效阻止医学误导信息的输出。结论指出,引导模型优先考虑逻辑合理性而非服从性,可有效缓解医疗领域被滥用的风险。
原文摘要 · Abstract (English)
Background: Large language models (LLMs) are trained to follow directions, but this introduces a vulnerability to blindly comply with user requests even if they generate wrong information. In medicine, this could accelerate the generation of misinformation that impacts human well-being. Objectives/Methods: We analyzed compliance to requests to generate misleading content about medications in settings where models know the request is illogical. We investigated whether in-context directions and instruction-tuning of LLMs to prioritize logical reasoning over compliance reduced misinformation risk. Results: While all frontier LLMs complied with misinformation requests, both prompt-based and parameter-based approaches can improve the detection of logic flaws in requests and prevent the dissemination of medical misinformation. Conclusion: Shifting LLMs to prioritize logic over compliance could reduce risks of exploitation for medical misinformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。