详细信息
TRELLIS: a reinforcement learning–enhanced typed-anchor-graph architecture for reasoning-oriented multi-hop retrieval-augmented generation ( EI收录)
文献类型:期刊文献
英文题名:TRELLIS: a reinforcement learning–enhanced typed-anchor-graph architecture for reasoning-oriented multi-hop retrieval-augmented generation
作者:Zheng, Yujia[1]; Gong, Hao[2]; Liu, Hongzhe[1]
第一作者:Zheng, Yujia
机构:[1] Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China; [2] School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China
第一机构:北京联合大学北京市信息服务工程重点实验室
通讯机构:[1]Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China|[11417103]北京联合大学北京市信息服务工程重点实验室;[11417]北京联合大学;
年份:2026
卷号:8
期号:12
外文期刊名:Engineering Research Express
收录:EI(收录号:20262520947995);Scopus(收录号:2-s2.0-105042150624)
语种:英文
外文关键词:Computational linguistics - Data mining - Graph structures - Information retrieval - Iterative methods - Multi agent systems - Optimization - Trajectories - Trees (mathematics)
摘要:Recent advances in large language models and dense retrievers have advanced retrieval-augmented generation (RAG), yet multi-hop reasoning tasks remain challenging. Existing approaches suffer from three limitations: (1) weak reasoning-oriented planning: they rely on hand-crafted or static decomposition strategies that overfit datasets and hop patterns; (2) suboptimal reasoning-driven retrieval: retrieval and query reformulation are only loosely coupled to the reasoning state, hindering fine-grained credit assignment to retrieval steps; and (3) insufficient reasoning-guided filtering: adaptive RAG frameworks mostly optimize final-answer metrics and give little supervision on which evidence to keep or discard. We propose typed retrieval-enhanced learning of latent inference structure TRELLIS, a trajectory-centric multi-hop RAG framework that views retrieval as constructing a tree-structured reasoning trajectory over a typed anchor graph. TRELLIS orchestrates three modules: a schema induction agent that induces executable anchor graphs, an evidence gating & extraction agent for calibrated evidence aggregation and answer extraction, and a relative-gain rewrite agent for query reformulation guided by relative retrieval gains. All decisions and counterfactuals are logged in a causal trace ledger and optimized via phase-coupled variant of group relative policy, a group-relative policy optimization method. Experiments on multi-hop question answering and retrieval benchmarks show that TRELLIS achieves competitive or improved accuracy over strong single-hop, iterative, and multi-agent RAG baselines on several benchmarks, while yielding structured, auditable retrieval trajectories. ? 2026 IOP Publishing Ltd. All rights, including for text and data mining, AI training, and similar technologies, are reserved. This article is available under the terms of the https://publishingsupport.iopscience.iop.org/iop-standard/v1.
参考文献:
正在载入数据...
