arXiv:2607.11564cs.CLcs.HC2026-07

用论文内容匹配个人文件夹,自动分类文献不需训练。

PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing

论文配图:PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing
图 1 · 摘自论文原文
  • 基于文件夹内已有论文内容做决策,而非仅看文件夹名。
  • 召回率从0.39提升至0.61,尤其在按期刊年份分的文件夹中提升显著。
  • 无需训练、低成本,适合个人文献管理场景使用。

研究人员在参考文献管理器中构建个性化文件夹层级,将新论文路由到对应文件夹。该任务不同于标准的层次文本分类:用户文件夹体系是私有的、动态演化的非正式分类体系,其含义可能是主题型、缩写型、会议型或流程型,常由已存论文定义。我们提出个性化层级论文路由(PHPR)问题:在不进行用户专属训练的前提下,将论文分配至用户特定层级。为此,我们设计PaperRouter-Agent,一个无需训练的LLM代理,通过分析文件夹成员论文内容来决定路由。该代理先缩小候选层级,检索文件夹特异性证据,通过检查成员论文验证匹配度,并融合历史用户拒收反馈。真实个人文献库的初步研究显示,整体Recall@1从0.39升至0.61,Recall@3从0.57升至0.83;在按期刊或年份定义的组织型文件夹中,单次方法召回率从0.09升至0.50。在公开的LaMP-2基准上,准确率从44.5%提升至51.5%(+9.0 macro-F1),且保持低计算成本。

原文摘要 · Abstract (English)

Researchers organize the papers they collect into personal folder hierarchies in reference managers, and route each new paper into the folder where it belongs. This task differs from standard hierarchical text classification. A user's folder hierarchy is not a fixed, shared taxonomy but a private and evolving folksonomy whose folder meanings may be topical, shorthand, venue-based, or process-oriented, and are often defined by the papers already stored inside them. We formalize this setting as personalized hierarchical paper routing (PHPR): assigning an incoming paper to folders in a user-specific hierarchy without per-user training. We propose PaperRouter-Agent, a training-free LLM agent that grounds routing decisions in folder members rather than folder names alone. The agent first narrows the candidate hierarchy, retrieves folder-specific evidence, verifies fit by inspecting member papers, and incorporates similarity-gated feedback from past user rejections. A formative study on real personal libraries shows that PaperRouter-Agent raises overall Recall@1 from 0.39 to 0.61 and Recall@3 from 0.57 to 0.83, with the largest gains on organizational folders defined by metadata such as venue or year, where single-shot methods collapses (Recall@1 0.09 to 0.50). On the public LaMP-2 benchmark, the same approach improves accuracy from 44.5% to 51.5% (+9.0 macro-F1) over a single-shot baseline, while remaining low-cost for practical use.

文献管理智能路由大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。