arXiv:2504.21489cs.CYcs.AI2025-04被引 3

打造真实场景下有效的AI检测评估框架,提升工具可信度与社会价值。

TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS

  • 基于真实案例与全球调研构建评估框架,强调可解释性与文化适配性。
  • 指出现有工具在多语言、跨文化场景中性能不足,需改进实测表现。
  • 适合开发者、政策制定者及标准组织用于设计透明、负责任的AI检测方案。

生成式AI与伪造合成媒体的泛滥正威胁全球信息生态,尤其影响全球南方地区。本报告指出,当前AI检测工具常因可解释性差、公平性不足、可及性低及语境相关性弱,在真实场景中表现不佳。为此,WITNESS提出真正创新且有效的AI检测评估基准(TRIED Benchmark),依据一线经验、典型欺诈案例与全球多方咨询,强调检测工具必须适应多元语言、文化和技术环境,才能实现真正创新与实用。报告为开发人员、政策制定者及标准机构提供具体指导,推动构建问责制、透明化、以用户为中心的检测解决方案,并将社会技术因素纳入未来AI标准与评估体系。采用TRIED基准有助于激发创新、维护公众信任、增强AI素养,助力构建更具韧性的全球信息可信生态。

原文摘要 · Abstract (English)

The proliferation of generative AI and deceptive synthetic media threatens the global information ecosystem, especially across the Global Majority. This report from WITNESS highlights the limitations of current AI detection tools, which often underperform in real-world scenarios due to challenges related to explainability, fairness, accessibility, and contextual relevance. In response, WITNESS introduces the Truly Innovative and Effective AI Detection (TRIED) Benchmark, a new framework for evaluating detection tools based on their real-world impact and capacity for innovation. Drawing on frontline experiences, deceptive AI cases, and global consultations, the report outlines how detection tools must evolve to become truly innovative and relevant by meeting diverse linguistic, cultural, and technological contexts. It offers practical guidance for developers, policy actors, and standards bodies to design accountable, transparent, and user-centered detection solutions, and incorporate sociotechnical considerations into future AI standards, procedures and evaluation frameworks. By adopting the TRIED Benchmark, stakeholders can drive innovation, safeguard public trust, strengthen AI literacy, and contribute to a more resilient global information credibility.

AI检测可信评估社会技术全球视野

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。