arXiv:2607.14673cs.AIcs.HC2026-07

为真实场景AI应用设计可定制的人工对齐评估流程。

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

论文配图:Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
图 1 · 摘自论文原文
  • 基于角色生成测试用例,结合上下文制定评分标准。
  • 人工标注与大模型评分联动,仅在一致时启用自动化。
  • 适合需合规、可审计的公共部门AI产品团队使用。

评估是真实世界AI应用部署的瓶颈:公开基准往往不匹配团队用户、场景或政策,而人工评审又难以扩展。针对我们在公共部门推进AI应用的经验,本文提出Kaleidoscope——一种集成的上下文化功能评估工作流,融合基于角色的测试用例生成、情境化评分标准与人工审核,实现可靠性保障的自动化评分。测试用例按特定应用的评分标准打分,人工标注提供可审查标签,大模型裁判仅在与标签一致率达到预设阈值时才参与评分。该流程具备可解释性、可迭代性,适用于实际产品团队。我们在四个组织用例中开展为期三周的试点,并在108个问答对上测试了四种领域的14项评估维度的自定义评分标准。结果表明,该框架支持端到端可靠的自动化评分。

原文摘要 · Abstract (English)

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

AI评估人机对齐公共部门可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。