为聊天机器人建立价值观审计框架,防范偏见与有害内容。
It is Time to Develop an Auditing Framework to Promote Value Aware Chatbots
- 提出基于社会价值的审计框架,统一评估聊天机器人表现。
- 测试GPT-3.5和GPT-4在搜索、编程、故事生成中存在违规范例。
- 呼吁学术界、政府、企业共建标准,推动技术向善。
2022年11月发布的ChatGPT开启了通用AI的新时代,使生成式AI工具普及化。这类聊天机器人具备解答作业、创作音乐与艺术等广泛能力,但其训练数据源自人类,不可避免继承错误与偏见,可能对特定群体造成伤害或加剧不平等。由于缺乏对社会价值的内在理解,它们生成的内容可能违背既有规范,如涉及儿童色情、事实错误或歧视性言论。本文主张,面对技术快速演进,计算机与数据科学家应迅速建立以价值观为核心的审计框架,包含社区共识的评估指标,用于持续监测不同聊天机器人及大模型的健康状态。我们提供一个简易审计模板,展示了在搜索引擎类任务、代码生成和故事生成中的初步审计结果,发现GPT-3.5与GPT-4存在符合与不符合现行法律价值的回应。尽管结果不令人意外,却凸显了建立可共享、标准化审计体系的紧迫性,以便学术界、政府与企业共同制定缓解策略。最后,论文提出若干基于价值观的技术改进建议。
原文摘要 · Abstract (English)
The launch of ChatGPT in November 2022 marked the beginning of a new era in AI, the availability of generative AI tools for everyone to use. ChatGPT and other similar chatbots boast a wide range of capabilities from answering student homework questions to creating music and art. Given the large amounts of human data chatbots are built on, it is inevitable that they will inherit human errors and biases. These biases have the potential to inflict significant harm or increase inequity on different subpopulations. Because chatbots do not have an inherent understanding of societal values, they may create new content that is contrary to established norms. Examples of concerning generated content includes child pornography, inaccurate facts, and discriminatory posts. In this position paper, we argue that the speed of advancement of this technology requires us, as computer and data scientists, to mobilize and develop a values-based auditing framework containing a community established standard set of measurements to monitor the health of different chatbots and LLMs. To support our argument, we use a simple audit template to share the results of basic audits we conduct that are focused on measuring potential bias in search engine style tasks, code generation, and story generation. We identify responses from GPT 3.5 and GPT 4 that are both consistent and not consistent with values derived from existing law. While the findings come as no surprise, they do underscore the urgency of developing a robust auditing framework for openly sharing results in a consistent way so that mitigation strategies can be developed by the academic community, government agencies, and companies when our values are not being adhered to. We conclude this paper with recommendations for value-based strategies for improving the technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。