arXiv:2410.10850cs.CLcs.AI2024-10被引 9

测试大模型在气候与心理领域的误导性问题上表现,发现回答正确但仍有伦理隐患。

On the Reliability of Large Language Models to Misinformed and Demographically-Informed Prompts

  • 用真假题测试模型判断力,闭合问题回答准确率高。
  • 专家评估发现存在隐私风险和引导专业服务缺失。
  • 适合关注AI伦理与医疗辅助的从业者参考。

我们研究了大型语言模型(LLM)驱动聊天机器人在气候变化与心理健康领域应对误导性提示和含人口统计信息问题时的表现。通过定量与定性方法,评估其辨别陈述真实性、遵循事实的能力,以及回应中是否存在偏见或错误信息。定量分析使用真假问答显示,这些聊天机器人在闭合问题上可提供正确答案。然而,领域专家的定性反馈表明,仍存在隐私、伦理问题,且聊天机器人未能充分引导用户寻求专业帮助。结论认为,尽管这些聊天机器人前景广阔,但在敏感领域部署需谨慎,须加强伦理监管与精细化优化,以作为人类专业能力的有益补充而非独立解决方案。

原文摘要 · Abstract (English)

We investigate and observe the behaviour and performance of Large Language Model (LLM)-backed chatbots in addressing misinformed prompts and questions with demographic information within the domains of Climate Change and Mental Health. Through a combination of quantitative and qualitative methods, we assess the chatbots' ability to discern the veracity of statements, their adherence to facts, and the presence of bias or misinformation in their responses. Our quantitative analysis using True/False questions reveals that these chatbots can be relied on to give the right answers to these close-ended questions. However, the qualitative insights, gathered from domain experts, shows that there are still concerns regarding privacy, ethical implications, and the necessity for chatbots to direct users to professional services. We conclude that while these chatbots hold significant promise, their deployment in sensitive areas necessitates careful consideration, ethical oversight, and rigorous refinement to ensure they serve as a beneficial augmentation to human expertise rather than an autonomous solution.

大模型伦理风险医疗辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。