用论文-专利对构建基准,提升长文本专利自动生成质量
Pap2Pat: Benchmarking Outline-Guided Long-Text Patent Generation with Patent-Paper Pairs
- 以论文为技术说明,分块引导生成专利描述
- 模型能利用论文信息但细节不足,微调后语言更专利化但幻觉增多
- 适合专利自动化、LLM长文本生成研究者使用
大语言模型在处理长而复杂的科技文本方面仍具挑战,尤其在耗时费力的专利撰写中尚未发挥潜力。专利说明书平均占文档90%以上,但其自动生成研究较少。撰写专利时常依赖保密的发明报告(IR),限制了研究进展;而预印本论文常可作为替代。本文利用这一双重性,构建了包含1.8k个专利-论文对的开放真实基准PAP2PAT,用于专利撰写任务。针对复杂长文档生成问题,提出基于分块与大纲引导的生成方法,以论文作为技术说明。通过PAP2PAT评估及人工案例研究发现,模型虽能有效利用论文信息,但仍难以提供足够细节;微调可增强专利风格语言,但也导致更多幻觉。数据与代码已开源:https://github.com/boschresearch/Pap2Pat。
原文摘要 · Abstract (English)
Dealing with long and highly complex technical text is a challenge for Large Language Models (LLMs), which still have to unfold their potential in supporting expensive and timeintensive processes like patent drafting. Within patents, the description constitutes more than 90% of the document on average. Yet, its automatic generation remains understudied. When drafting patent applications, patent attorneys typically receive invention reports (IRs), which are usually confidential, hindering research on LLM-supported patent drafting. Often, prepublication research papers serve as IRs. We leverage this duality to build PAP2PAT, an open and realistic benchmark for patent drafting consisting of 1.8k patent-paper pairs describing the same inventions. To address the complex longdocument patent generation task, we propose chunk-based outline-guided generation using the research paper as invention specification. Our extensive evaluation using PAP2PAT and a human case study show that LLMs can effectively leverage information from the paper, but still struggle to provide the necessary level of detail. Fine-tuning leads to more patent-style language, but also to more hallucination. We release our data and code https://github.com/boschresearch/Pap2Pat.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。