arXiv:2508.16753cs.CL2025-08AAAI被引 1

GAICo统一评估生成式AI的多模态输出,让对比更高效可靠。

GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs

  • 提供统一框架,支持文本、结构化数据和多媒体的评估
  • 可快速实现多模型对比、可视化与报告生成,提升开发效率
  • 开源工具已获超1.6万次下载,适合研发与部署阶段使用

生成式AI在多样化高风险领域快速普及,亟需稳健且可复现的评估方法。然而,实践者常依赖非标准化脚本,现有指标对特定结构化输出(如自动计划、时间序列)或跨模态(文本、音频、图像)整体比较不适用,导致评估碎片化,阻碍系统开发。为此,我们提出GAICo(Generative AI Comparator):一个已部署的开源Python库,旨在简化并标准化生成式AI输出的对比。GAICo提供统一、可扩展的框架,支持针对无结构文本、特定结构数据格式及多媒体(图像、音频)的参考基准度量。其架构包含高层API,可实现从多模型对比到可视化与报告生成的端到端分析,同时支持直接调用度量以实现细粒度控制。我们通过一个详细案例研究,展示了其在评估和调试复杂多模态AI旅行助手流水线中的应用。GAICo助力研究人员与开发者高效评估系统性能,实现评估可复现,提升开发速度,最终构建更可信的AI系统,推动安全高效的AI部署。自2025年6月在PyPI发布以来,截至2025年12月,该工具累计下载量超过16,000次,体现了社区日益增长的兴趣。

原文摘要 · Abstract (English)

The rapid proliferation of Generative AI (GenAI) into diverse, high-stakes domains necessitates robust and reproducible evaluation methods. However, practitioners often resort to ad-hoc, non-standardized scripts, as common metrics are often unsuitable for specialized, structured outputs (e.g., automated plans, time-series) or holistic comparison across modalities (e.g., text, audio, and image). This fragmentation hinders comparability and slows AI system development. To address this challenge, we present GAICo (Generative AI Comparator): a deployed, open-source Python library that streamlines and standardizes GenAI output comparison. GAICo provides a unified, extensible framework supporting a comprehensive suite of reference-based metrics for unstructured text, specialized structured data formats, and multimedia (images, audio). Its architecture features a high-level API for rapid, end-to-end analysis, from multi-model comparison to visualization and reporting, alongside direct metric access for granular control. We demonstrate GAICo's utility through a detailed case study evaluating and debugging complex, multi-modal AI Travel Assistant pipelines. GAICo empowers AI researchers and developers to efficiently assess system performance, make evaluation reproducible, improve development velocity, and ultimately build more trustworthy AI systems, aligning with the goal of moving faster and safer in AI deployment. Since its release on PyPI in Jun 2025, the tool has been downloaded over 16K times, across versions, by Dec 2025, demonstrating growing community interest.

生成式AI多模态评估开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。