arXiv:2512.15791cs.CYcs.AI2025-12中稿 · publication in AI …

评估AI伦理工具在语言模型中的实用性,发现其难以识别特定语言的潜在风险。

Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study

  • 通过文献梳理与开发者访谈,系统评估4种主流伦理工具
  • 工具能引导通用伦理思考,但无法捕捉葡语特有表达的风险
  • 适合关注AI责任设计的开发团队与政策制定者参考

在人工智能领域,语言模型因能够通过文本生成模拟人类对话而广泛应用。鉴于其对社会的重大影响,负责任地开发和部署这些模型至关重要,需关注其潜在负面影响。近年来,针对这一需求,涌现出大量人工智能伦理工具(AIETs),旨在帮助开发者、企业、政府等利益相关方建立信任、透明度与责任感。然而,多数工具缺乏清晰文档、使用示例及实际有效性验证。本文提出一种评估方法,基于对213项AIETs的文献综述,筛选出4个代表性工具:Model Cards、ALTAI、FactSheets与Harms Modeling。研究将这些工具应用于葡萄牙语语言模型,并对35小时开发者访谈进行分析,从开发者视角评估其在识别模型伦理问题方面的表现。结果表明,这些工具可作为通用伦理考量的指引,但未能覆盖语言模型的独特特征,如习语表达;且未有效识别葡语模型可能带来的负面效应。

原文摘要 · Abstract (English)

In Artificial Intelligence (AI), language models have gained significant importance due to the widespread adoption of systems capable of simulating realistic conversations with humans through text generation. Because of their impact on society, developing and deploying these language models must be done responsibly, with attention to their negative impacts and possible harms. In this scenario, the number of AI Ethics Tools (AIETs) publications has recently increased. These AIETs are designed to help developers, companies, governments, and other stakeholders establish trust, transparency, and responsibility with their technologies by bringing accepted values to guide AI's design, development, and use stages. However, many AIETs lack good documentation, examples of use, and proof of their effectiveness in practice. This paper presents a methodology for evaluating AIETs in language models. Our approach involved an extensive literature survey on 213 AIETs, and after applying inclusion and exclusion criteria, we selected four AIETs: Model Cards, ALTAI, FactSheets, and Harms Modeling. For evaluation, we applied AIETs to language models developed for the Portuguese language, conducting 35 hours of interviews with their developers. The evaluation considered the developers' perspective on the AIETs' use and quality in helping to identify ethical considerations about their model. The results suggest that the applied AIETs serve as a guide for formulating general ethical considerations about language models. However, we note that they do not address unique aspects of these models, such as idiomatic expressions. Additionally, these AIETs did not help to identify potential negative impacts of models for the Portuguese language.

AI伦理语言模型开发者视角工具评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。