为通用AI设计负责任的系统,需解决高自由度带来的幻觉与偏见问题。
Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward
- 提出C2V2框架:控制、一致、价值、真实,应对通用AI风险
- 通用AI输出自由度高,导致幻觉、偏见更严重且难控
- 适合关注AI伦理、安全与系统设计的研究者与开发者
现代通用人工智能系统(如大语言模型和视觉模型)可完成文本撰写、代码生成与调试、数据库查询、多语言翻译等多种任务,广受行业欢迎。然而,其输出存在幻觉、毒性与刻板印象等风险,难以信任。本文基于八项公认的负责任AI(RAI)原则(公平性、隐私、可解释性、鲁棒性、安全性、真实性、治理、可持续性),分析通用AI的风险与脆弱性,并对比传统任务专用系统(风险更低且易缓解)。原因在于通用AI输出具有非确定性的高自由度(DoFo),而传统系统则为确定性低或常数级自由度。为此,本文提出C2V2(控制、一致性、价值、真实性)理想要求,并评估当前AI对齐、检索增强生成、推理增强等方法在达成这些目标上的表现。我们主张,未来通用AI应通过形式化建模特定应用/领域的RAI需求,结合系统化设计整合多种技术,以满足C2V2维度的要求。
原文摘要 · Abstract (English)
Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vulnerabilities of modern general-purpose AI along eight widely accepted responsible AI (RAI) principles (fairness, privacy, explainability, robustness, safety, truthfulness, governance, and sustainability) and compare how they are non-existent or less severe and easily mitigable in traditional task-specific counterparts. We argue that this is due to the non-deterministically high Degree of Freedom in output (DoFo) of general-purpose AI (unlike the deterministically constant or low DoFo of traditional task-specific AI systems), and there is a need to rethink our approach to RAI for general-purpose AI. Following this, we derive C2V2 (Control, Consistency, Value, Veracity) desiderata to meet the RAI requirements for future general-purpose AI systems, and discuss how recent efforts in AI alignment, retrieval-augmented generation, reasoning enhancements, etc. fare along one or more of the desiderata. We believe that the goal of developing responsible general-purpose AI can be achieved by formally modeling application- or domain-dependent RAI requirements along C2V2 dimensions, and taking a system design approach to suitably combine various techniques to meet the desiderata.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。