23个大模型在责任价值上更接近从业者而非普通人,但言行不一。
Are We Aligned? A Preliminary Investigation of the Alignment of Responsible AI Values between LLMs and Human Judgment
- 对比人类群体评估23个大模型的责任价值偏好
- 模型对公平、隐私等价值认同度高,但优先级决策不一致
- 适合关注AI伦理与软件工程结合的研究者
大型语言模型(LLMs)被越来越多地用于需求获取、设计与评估等软件工程任务,引发其与人类在负责任AI价值观上是否对齐的关切。本研究考察了23个LLMs在四个任务中的表现:(T1)选择关键负责任AI价值,(T2)在具体情境中评估其重要性,(T3)解决相互冲突的价值权衡,(T4)优先排序体现这些价值的软件需求。结果显示,LLMs整体上与AI从业者比美国代表性人群更一致,强调公平、隐私、透明、安全和问责;但在宣称的价值(任务1-3)与实际需求优先级(任务4)之间存在不一致,暴露出陈述与行为间的忠实度差距。这表明在无监督情况下依赖LLMs进行需求工程存在实际风险,亟需建立系统性的基准测试、解释与监控机制以保障价值对齐。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly employed in software engineering tasks such as requirements elicitation, design, and evaluation, raising critical questions regarding their alignment with human judgments on responsible AI values. This study investigates how closely LLMs' value preferences align with those of two human groups: a US-representative sample and AI practitioners. We evaluate 23 LLMs across four tasks: (T1) selecting key responsible AI values, (T2) rating their importance in specific contexts, (T3) resolving trade-offs between competing values, and (T4) prioritizing software requirements that embody those values. The results show that LLMs generally align more closely with AI practitioners than with the US-representative sample, emphasizing fairness, privacy, transparency, safety, and accountability. However, inconsistencies appear between the values that LLMs claim to uphold (Tasks 1-3) and the way they prioritize requirements (Task 4), revealing gaps in faithfulness between stated and applied behavior. These findings highlight the practical risk of relying on LLMs in requirements engineering without human oversight and motivate the need for systematic approaches to benchmark, interpret, and monitor value alignment in AI-assisted software development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。