大模型评分在表格识别中不可靠,无法有效指导迭代优化。
LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

- 用大模型做评分器评估表格生成结果,但得分常相同、排名不一致。
- 多次迭代后模型表现未提升,评分器仅部分恢复改进结果。
- 需结构验证信号而非仅依赖评分,适合研究闭环生成的学者参考。
LLM-as-a-judge广泛用于闭环再生中的反馈与选择信号,但其有效性尚未充分验证。本文在表格识别任务中,利用确定性TEDS评估作为受控测试平台,基于FinTabNet和OmniDocBench展开研究。发现三方面问题:第一,评分信号弱——得分频繁相同,排名不可复现;无论采用候选选择还是保守分数阈值接受策略,均未优于首次输出;迭代虽产生更优候选,但评分器仅在一项数据集部分恢复,另一项完全未能识别。第二,即使无具体评分反馈,仍出现严重性能下降,表明非约束重生成下目标保持失败是关键机制。第三,加入结构保持指令显著降低严重损失率(在FinTabNet上明显,在OmniDocBench上趋势向好),但在保留评分反馈的探索性2×2分析中,该保护效果未稳定存在。结论并非否定大模型作为评估者的价值,而是指出当前无参考评分信号过于薄弱且不稳定,无法支撑此场景下的候选选择,仅凭评估证据不足以证明闭环优化的有效性。迭代优化至少需要能确定检测结构变化的验证信号,而非仅依赖评分。
原文摘要 · Abstract (English)
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently validated. We study it in table recognition, where deterministic TEDS evaluation provides a controlled testbed, using FinTabNet and OmniDocBench. Three findings emerge. First, judge signals were weak on both datasets: scores frequently tied, rankings were not reproducible, and no tested judge score policy, whether selecting among candidates or accepting revisions under a conservative score margin, improved on the first output on both datasets. Iteration produced better candidates, but the judge recovered them at most partially on one dataset and not at all on the other. Second, severe losses occurred even without specific judge feedback, supporting target-preservation failure under unconstrained regeneration as a proximate mechanism. Third, a structure-preserving instruction reduced the severe-loss rate, significantly on FinTabNet and directionally on OmniDocBench, but produced no improvement, and in an exploratory 2x2 analysis this protection was not stably observed when judge feedback was retained. These results do not dispute the value of LLMs as evaluators, but show that the tested reference-free judge signals were too weak and unstable to drive candidate selection in this setup, and that evaluation-style evidence alone was insufficient to establish closed-loop optimization utility. Iterative refinement requires, at minimum, a verification signal that deterministically detects structural change, rather than judge scores alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。