arXiv:2504.16778cs.CLcs.AI2025-04被引 7

为真实场景中的生成式AI设计动态评估框架

Evaluation Framework for AI Systems in "the Wild"

  • 提出融合多样、动态输入的综合评估方法
  • 强调持续监测与人机结合的评估机制
  • 适合关注伦理与社会影响的开发者和政策制定者

生成式AI已广泛应用于各行业,但现有评估方法仍依赖固定基准和数据集,难以反映实际应用表现,导致实验室结果与现实效果脱节。本文提出一套面向真实场景的生成式AI评估框架,强调使用多样化、不断变化的输入,并采用全面、动态、持续的评估方式。该框架支持实践者设计能反映实时能力的评估方案,也为政策制定者提供基于社会影响而非固定性能指标或参数规模的政策建议。倡导整合性能、公平性与伦理的综合性评估体系,结合人类与自动化评估,保持透明以建立利益相关方信任。实施这些策略可确保生成式AI不仅技术先进,且具备伦理责任与实际影响力。

原文摘要 · Abstract (English)

Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on benchmarks and fixed datasets, frequently failing to reflect real-world performance, which creates a gap between lab-tested outcomes and practical applications. This white paper proposes a comprehensive framework for how we should evaluate real-world GenAI systems, emphasizing diverse, evolving inputs and holistic, dynamic, and ongoing assessment approaches. The paper offers guidance for practitioners on how to design evaluation methods that accurately reflect real-time capabilities, and provides policymakers with recommendations for crafting GenAI policies focused on societal impacts, rather than fixed performance numbers or parameter sizes. We advocate for holistic frameworks that integrate performance, fairness, and ethics and the use of continuous, outcome-oriented methods that combine human and automated assessments while also being transparent to foster trust among stakeholders. Implementing these strategies ensures GenAI models are not only technically proficient but also ethically responsible and impactful.

生成式AI评估框架伦理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。