arXiv:2608.11381cs.AIcs.LG2026-08

用专家代理+强化学习提升房地产金融分析的准确性。

From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate

论文配图:From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
图 1 · 摘自论文原文
  • 拆分任务为8个专业代理,提升数值计算精度
  • 强化学习使判断类任务得分提高14.2点
  • 新模型对未见公司和监管框架有良好泛化能力

我们研究本地化数值操作与综合判断在金融分析中是否受益于相同的LLM专业化形式。Larix将16个镜头的欧洲上市房地产分析框架映射为8个对齐专业的代理;在保持模型、证据源、任务指令、输出格式和评分标准不变的前提下,对比了前沿LLM在整体提示与专家分解提示下的表现。在覆盖7种监管框架的19家公司的测试中,分解使数值任务总分提升15.8个百分点,但对判断任务效果不显著,甚至可能降低,该模式在四次固定模板调度中均稳定存在;单一代理使用完整框架无法复现数值提升。随后,使用任务对齐结构奖励对Qwen3.5-9B进行后训练GRPO,使开发集得分提升12.0点,判断总分提升14.2点,所有四个子任务均获益;在未见公司上总体提升15.2点(契约压力任务提升40.4点),在未见监管框架上提升4.3点,且在所有三组防记忆化划分中均有正向迁移。提示级分解改善模块化数值执行,而针对性参数调整则提升整合性金融判断。

原文摘要 · Abstract (English)

We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points but does not reliably improve, and can reduce, performance on judgment tasks, a pattern stable across four frozen-template dispatches; a single-agent control given the complete framework does not reproduce the numerical gain. Post-training Qwen3.5-9B with GRPO using task-aligned structured rewards then raises the development-split score by 12.0 points and the judgment aggregate by 14.2 points, with gains on all four sub-ceiling tasks; the gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress) and to unseen regulatory wrappers (+4.3), with positive transfer on all three anti-memorization splits. Prompt-level decomposition thus improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment.

金融分析LLM专家强化学习房地产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。