arXiv:2511.05496cs.IRcs.AI2025-11

用大模型构建可定制的文档评估流程,提升评估效率与可追溯性。

DOCUEVAL: An LLM-based AI Engineering Tool for Building Customisable Document Evaluation Workflows

  • 基于大模型设计可自定义的文档评估工作流,支持角色与标准灵活配置。
  • 通过完整日志与配置管理,实现评估结果的可比性与可复现性。
  • 适用于需要高可靠性评估的科研评审、内容质检等场景。

基础模型(如大语言模型)有望简化评估流程并提升性能,但实际应用仍面临可定制性、准确性与可扩展性的挑战。本文提出 DOCUEVAL,一个用于构建可定制文档评估工作流的 AI 工程工具。DOCUEVAL 支持高级文档处理与可定制工作流设计,允许用户定义基于理论的评审角色、指定评估标准、实验不同推理策略并选择评估风格。为确保可追溯性,DOCUEVAL 提供每轮运行的完整日志、来源标注与配置管理,支持不同设置下的结果系统性对比。通过整合这些能力,DOCUEVAL 直接应对核心软件工程问题:如何判断评估者是否“足够好”以部署,以及如何实证比较不同评估策略。我们通过真实学术同行评审案例验证了 DOCUEVAL 的有效性,展示了其在评估者工程化与可扩展、可靠文档评估中的价值。

原文摘要 · Abstract (English)

Foundation models, such as large language models (LLMs), have the potential to streamline evaluation workflows and improve their performance. However, practical adoption faces challenges, such as customisability, accuracy, and scalability. In this paper, we present DOCUEVAL, an AI engineering tool for building customisable DOCUment EVALuation workflows. DOCUEVAL supports advanced document processing and customisable workflow design which allow users to define theory-grounded reviewer roles, specify evaluation criteria, experiment with different reasoning strategies and choose the assessment style. To ensure traceability, DOCUEVAL provides comprehensive logging of every run, along with source attribution and configuration management, allowing systematic comparison of results across alternative setups. By integrating these capabilities, DOCUEVAL directly addresses core software engineering challenges, including how to determine whether evaluators are "good enough" for deployment and how to empirically compare different evaluation strategies. We demonstrate the usefulness of DOCUEVAL through a real-world academic peer review case, showing how DOCUEVAL enables both the engineering of evaluators and scalable, reliable document evaluation.

文档评估LLM工程可追溯性工作流设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。