arXiv:2512.08449cs.AI2025-12

用理论变革框架让AI系统真正产生社会影响

From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change

  • 将变革理论五阶段映射到AI架构层,实现价值对齐
  • 通过多目标优化和因果图减少幻觉,提升公平性
  • 适合关注伦理与社会影响的AI开发者和政策制定者

本文提出影响驱动型人工智能框架(IDAIF),将变革理论(ToC)的五个阶段(输入-活动-产出-结果-影响)与AI架构层(数据层-管道层-推理层-智能体层-规范层)系统性对应。针对医疗、金融和公共政策等高风险领域中AI行为与人类价值观脱节的问题,IDAIF引入多目标帕累托优化实现价值对齐,采用分层多智能体编排达成结果,利用因果有向无环图(DAG)抑制幻觉,并结合基于人类反馈的强化学习(RLHF)进行对抗性去偏以保障公平。框架还设有保障层,通过守护架构处理假设失效问题。三个案例研究验证了其在医疗、网络安全和软件工程中的应用效果。该框架标志着从模型中心转向影响中心的范式转变,为构建可信、有益的社会化AI系统提供了可落地的架构模式。

原文摘要 · Abstract (English)

This paper introduces the Impact-Driven AI Framework (IDAIF), a novel architectural methodology that integrates Theory of Change (ToC) principles with modern artificial intelligence system design. As AI systems increasingly influence high-stakes domains including healthcare, finance, and public policy, the alignment problem--ensuring AI behavior corresponds with human values and intentions--has become critical. Current approaches predominantly optimize technical performance metrics while neglecting the sociotechnical dimensions of AI deployment. IDAIF addresses this gap by establishing a systematic mapping between ToC's five-stage model (Inputs-Activities-Outputs-Outcomes-Impact) and corresponding AI architectural layers (Data Layer-Pipeline Layer-Inference Layer-Agentic Layer-Normative Layer). Each layer incorporates rigorous theoretical foundations: multi-objective Pareto optimization for value alignment, hierarchical multi-agent orchestration for outcome achievement, causal directed acyclic graphs (DAGs) for hallucination mitigation, and adversarial debiasing with Reinforcement Learning from Human Feedback (RLHF) for fairness assurance. We provide formal mathematical formulations for each component and introduce an Assurance Layer that manages assumption failures through guardian architectures. Three case studies demonstrate IDAIF application across healthcare, cybersecurity, and software engineering domains. This framework represents a paradigm shift from model-centric to impact-centric AI development, providing engineers with concrete architectural patterns for building ethical, trustworthy, and socially beneficial AI systems.

AI伦理架构设计变革理论可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。