arXiv:2512.03068cs.CYcs.AI2025-12被引 2

通过人与大模型协作,提前预测AI偏见引发的伤害。

Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation

  • 构建偏见到伤害的映射路径,识别潜在风险。
  • 在医疗和招聘领域验证,发现特定偏见导致的伤害模式。
  • 适合AI设计、治理及伦理审查人员使用。

人工智能在关键决策领域的广泛应用暴露了其可能造成的严重伤害,这些伤害往往源于整个生命周期中的偏见。现有框架多孤立地记录偏见或伤害,极少系统关联具体偏见类型与对应伤害,尤其缺乏对真实社会技术情境的考量。现有技术修复措施通常在系统开发或部署后才应用,难以实现预防性干预。本文提出ECHO框架,通过系统化映射不同偏见类型与伤害结果,在多元利益相关者和领域背景下实现主动式AI伤害预判。ECHO采用模块化流程,包括利益相关者识别、基于情景的偏见系统展示,以及人类与大语言模型双源伤害标注,并整合于伦理矩阵以结构化解读。该以人为本的方法可实现早期识别偏见到伤害的传导路径,从源头指导AI设计与治理决策。我们在疾病诊断与招聘两个高风险领域验证了ECHO,揭示了领域特异性的偏见-伤害模式,证明其在支持前瞻性AI治理方面的潜力。

原文摘要 · Abstract (English)

The growing influence of Artificial Intelligence (AI) systems on decision-making in critical domains has exposed their potential to cause significant harms, often rooted in biases embedded across the AI lifecycle. While existing frameworks and taxonomies document bias or harms in isolation, they rarely establish systematic links between specific bias types and the harms they cause, particularly within real-world sociotechnical contexts. Technical fixes proposed to address AI biases are ill-equipped to address them and are typically applied after a system has been developed or deployed, offering limited preventive value. We propose ECHO, a novel framework for proactive AI harm anticipation through the systematic mapping of AI bias types to harm outcomes across diverse stakeholder and domain contexts. ECHO follows a modular workflow encompassing stakeholder identification, vignette-based presentation of biased AI systems, and dual (human-LLM) harm annotation, integrated within ethical matrices for structured interpretation. This human-centered approach enables early-stage detection of bias-to-harm pathways, guiding AI design and governance decisions from the outset. We validate ECHO in two high-stakes domains (disease diagnosis and hiring), revealing domain-specific, bias-to-harm patterns and demonstrating ECHO's potential to support anticipatory governance of AI systems

AI伦理偏见检测风险预判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。