CIRCLE让AI落地效果可衡量,帮决策者看清真实影响。
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
- 六阶段框架把实际场景问题转为可测指标
- 结合实地测试与长期追踪,产出跨地点可比证据
- 适合关注AI落地实效的管理者与治理者
本文提出CIRCLE,一个基于生命周期的六阶段框架,弥合模型性能指标与AI实际部署效果之间的现实差距。现有MLOps和模型评测体系虽能揭示系统稳定性和模型能力,却无法为非技术决策者提供系统在真实环境中行为及其对组织长期影响的系统性证据。CIRCLE通过形式化将栈外利益相关方关切转化为可测量信号,操作化实施TEVV(测试、评估、验证、验证)中的验证阶段。不同于局部化的参与式设计或事后算法审计,CIRCLE提供结构化、前瞻性的协议,连接情境敏感的定性洞察与可扩展的定量指标。通过整合实地测试、红队演练和纵向研究形成协同流程,生成可跨站点比较但又保留本地情境敏感性的系统性知识,从而支持基于实际下游影响的治理,而非仅依赖理论能力。
原文摘要 · Abstract (English)
This paper proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and AI's materialized outcomes in deployment. Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insights into system stability and model capabilities, but they do not provide decision-makers outside the AI stack with systematic evidence of how these systems actually behave in real-world contexts or affect their organizations over time. CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by formalizing the translation of stakeholder concerns outside the stack into measurable signals. Unlike participatory design, which often remains localized, or algorithmic audits, which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context-sensitive qualitative insights to scalable quantitative metrics. By integrating methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline, CIRCLE produces systematic knowledge: evidence that is comparable across sites yet sensitive to local context. This, in turn, can enable governance based on materialized downstream effects rather than theoretical capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。