用AI模拟大法官提问,帮律师练口才。
AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments
- 构建双层评估框架,兼顾真实感与教学价值
- 模拟提问能准确覆盖核心法律议题,召回率高
- 但问题类型单一、过于迎合,需警惕评估盲区
在口头辩论中,法官会针对事实依据、法律主张和论点强度向律师提问。为应对此类质询,法学院和执业律师普遍采用模拟法庭(moot courts)进行演练。本文基于美国最高法院口头辩论记录数据集,探究AI模型是否可有效模拟特定大法官的提问风格。由于单次提问无唯一正确答案,评估应关注其是否具备预见实质性法律问题、识别逻辑漏洞及保持适度对抗语气等特质。我们提出一个两层评估框架,通过互补代理指标衡量模拟提问的真实感与教学实用性。构建并测试了基于提示(prompt-based)与自主代理(agentic)的两种模拟系统。结果显示,人类标注者普遍认为模拟问题具有现实感,且对真实法律议题的召回率较高。然而,模型仍存在提问类型贫乏、倾向讨好等显著缺陷,这些不足在简单评估方法下难以察觉。
原文摘要 · Abstract (English)
In oral arguments, judges probe attorneys with questions about the factual record, legal claims, and the strength of their arguments. To prepare for this questioning, both law schools and practicing attorneys rely on moot courts: practice simulations of appellate hearings. Leveraging a dataset of U.S. Supreme Court oral argument transcripts, we examine whether AI models can effectively simulate justice-specific questioning for moot court-style training. Evaluating oral argument simulation is challenging because there is no single correct question for any given turn. Instead, effective questioning should reflect a combination of desirable qualities, such as anticipating substantive legal issues, detecting logical weaknesses, and maintaining an appropriately adversarial tone. We introduce a two-layer evaluation framework that assesses both the realism and pedagogical usefulness of simulated questions using complementary proxy metrics. We construct and evaluate both prompt-based and agentic oral argument simulators. We find that simulated questions are often perceived as realistic by human annotators and achieve high recall of ground truth substantive legal issues. However, models still face substantial shortcomings, including low diversity in question types and sycophancy. Importantly, these shortcomings would remain undetected under naive evaluation approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。