用代理理论框架解决组织用大模型时的对齐难题
Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- 基于代理理论构建大模型对齐框架,应对信息不对称
- 提出组织采用大模型的分阶段对齐分析方法
- 首个系统梳理大模型对齐问题与解决方案的空间
将大语言模型(LLM)引入组织可能彻底改变我们的生活与工作方式。然而,这些模型可能生成无关、歧视性或有害内容。这一人工智能对齐问题通常源于采用过程中的规范缺失,由于模型的黑箱特性,委托方难以察觉。尽管多个研究领域探讨过AI对齐,但均未关注组织采用者与黑箱模型代理间的信息不对称,也未考虑组织级采用流程。为此,本文提出LLM ATLAS(基于代理理论的大模型对齐策略),一个以代理(合同)理论为基础的概念框架,用于缓解组织采用过程中的对齐问题。通过结合组织采用阶段与代理理论进行概念文献分析,本研究实现了:(1) 针对组织级大模型采用的对齐方法扩展文献分析流程;(2) 构建首个大模型对齐问题-解决方案空间。
原文摘要 · Abstract (English)
Adopting Large language models (LLMs) in organizations potentially revolutionizes our lives and work. However, they can generate off-topic, discriminating, or harmful content. This AI alignment problem often stems from misspecifications during the LLM adoption, unnoticed by the principal due to the LLM's black-box nature. While various research disciplines investigated AI alignment, they neither address the information asymmetries between organizational adopters and black-box LLM agents nor consider organizational AI adoption processes. Therefore, we propose LLM ATLAS (LLM Agency Theory-Led Alignment Strategy) a conceptual framework grounded in agency (contract) theory, to mitigate alignment problems during organizational LLM adoption. We conduct a conceptual literature analysis using the organizational LLM adoption phases and the agency theory as concepts. Our approach results in (1) providing an extended literature analysis process specific to AI alignment methods during organizational LLM adoption and (2) providing a first LLM alignment problem-solution space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。