arXiv:2605.10593cs.AIcs.CL2026-05中稿 · IJCAI

LLARS让领域专家与开发者协同构建大模型应用,全程无缝衔接。

LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation

论文配图:LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation
图 1 · 摘自论文原文
  • 支持实时协作撰写提示词,带版本控制和即时测试
  • 可批量生成多组合提示+模型+数据输出,并控制成本
  • 人机联合评估输出,一键转化完成批次为评估任务

我们展示了LLARS(大语言模型辅助研究系统),一个开源平台,弥合了领域专家与开发者在构建基于大模型系统时的鸿沟。该平台将三个紧密集成的模块整合为端到端流程:协作式提示工程支持实时共同创作,具备版本控制和即时大模型测试;批量生成支持用户选择的提示词×模型×数据组合,实现可配置输出并控制成本;混合评估通过人类与大模型评估者联合使用多种评估方法,提供实时一致率指标与溯源分析,以确定特定应用场景下的最优模型-提示组合。新提示和模型可自动用于批量生成,已完成批次可一键转化为评估场景。对六位领域专家和三位在线心理咨询开发者的访谈表明,LLARS操作直观,显著节省时间,使跨学科协作更加顺畅。

原文摘要 · Abstract (English)

We demonstrate LLARS (LLM Assisted Research System), an open-source platform that bridges the gap between domain experts and developers for building LLM-based systems. It integrates three tightly connected modules into an end-to-end pipeline: Collaborative Prompt Engineering for real-time co-authoring with version control and instant LLM testing, Batch Generation for configurable output production across user-selected prompts $\times$ models $\times$ data with cost control, and Hybrid Evaluation where human and LLM evaluators jointly assess outputs through diverse assessment methods, with live agreement metrics and provenance analysis to identify the best model-prompt combination for a given use case. New prompts and models are automatically available for batch generation and completed batches can be turned into evaluation scenarios with a single click. Interviews with six domain experts and three developers in online counselling confirmed that LLARS feels intuitive, saves considerable time by keeping everything in one place and makes interdisciplinary collaboration seamless.

提示工程人机协作评估系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。