arXiv:2509.04505cs.AIcs.CY2025-09被引 4

评估大模型在建筑项目管理中的伦理决策能力,发现其需人工监督。

The Ethical Compass of the Machine: Evaluating Large Language Models for Decision Support in Construction Project Management

  • 用新开发的伦理评估清单测试两模型在12个真实场景表现
  • 模型在合规类任务表现尚可,但缺乏上下文理解与透明推理
  • 专家强烈反对完全依赖AI做伦理判断,主张人机协同

人工智能在建筑项目管理中的应用加速,大语言模型正成为易用的决策支持工具。本研究旨在评估大模型在高风险、伦理敏感的建筑决策场景中的伦理可行性与可靠性。采用混合方法:用新型伦理决策支持评估清单(EDSAC)对两个主流大模型进行定量测试,覆盖12个真实伦理场景;同时对12位行业专家进行半结构化访谈,获取专业认知。结果表明,大模型在法律合规等结构化领域表现尚可,但在处理上下文细节、责任追溯和推理透明性方面存在显著缺陷。受访者普遍对自主使用AI进行伦理判断表示担忧,强烈建议实施严格的“人在回路”监管机制。本研究为首个在建筑领域实证检验大模型伦理推理的研究,提出可复用的EDSAC框架,并强调当前大模型应作为辅助工具而非独立伦理主体。

原文摘要 · Abstract (English)

The integration of Artificial Intelligence (AI) into construction project management (CPM) is accelerating, with Large Language Models (LLMs) emerging as accessible decision-support tools. This study aims to critically evaluate the ethical viability and reliability of LLMs when applied to the ethically sensitive, high-risk decision-making contexts inherent in CPM. A mixed-methods research design was employed, involving the quantitative performance testing of two leading LLMs against twelve real-world ethical scenarios using a novel Ethical Decision Support Assessment Checklist (EDSAC), and qualitative analysis of semi-structured interviews with 12 industry experts to capture professional perceptions. The findings reveal that while LLMs demonstrate adequate performance in structured domains such as legal compliance, they exhibit significant deficiencies in handling contextual nuance, ensuring accountability, and providing transparent reasoning. Stakeholders expressed considerable reservations regarding the autonomous use of AI for ethical judgments, strongly advocating for robust human-in-the-loop oversight. To our knowledge, this is one of the first studies to empirically test the ethical reasoning of LLMs within the construction domain. It introduces the EDSAC framework as a replicable methodology and provides actionable recommendations, emphasising that LLMs are currently best positioned as decision-support aids rather than autonomous ethical agents.

大模型伦理评估建筑管理人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。