登录    注册    忘记密码

详细信息

TRELLIS: a reinforcement learning–enhanced typed-anchor-graph architecture for reasoning-oriented multi-hop retrieval-augmented generation  ( EI收录)  

文献类型:期刊文献

英文题名:TRELLIS: a reinforcement learning–enhanced typed-anchor-graph architecture for reasoning-oriented multi-hop retrieval-augmented generation

作者:Zheng, Yujia[1]; Gong, Hao[2]; Liu, Hongzhe[1]

第一作者:Zheng, Yujia

机构:[1] Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China; [2] School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China

第一机构:北京联合大学北京市信息服务工程重点实验室

通讯机构:[1]Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing, China|[11417103]北京联合大学北京市信息服务工程重点实验室;[11417]北京联合大学;

年份:2026

卷号:8

期号:12

外文期刊名:Engineering Research Express

收录:EI(收录号:20262520947995);Scopus(收录号:2-s2.0-105042150624)

语种:英文

外文关键词:Computational linguistics - Data mining - Graph structures - Information retrieval - Iterative methods - Multi agent systems - Optimization - Trajectories - Trees (mathematics)

摘要:Recent advances in large language models and dense retrievers have advanced retrieval-augmented generation (RAG), yet multi-hop reasoning tasks remain challenging. Existing approaches suffer from three limitations: (1) weak reasoning-oriented planning: they rely on hand-crafted or static decomposition strategies that overfit datasets and hop patterns; (2) suboptimal reasoning-driven retrieval: retrieval and query reformulation are only loosely coupled to the reasoning state, hindering fine-grained credit assignment to retrieval steps; and (3) insufficient reasoning-guided filtering: adaptive RAG frameworks mostly optimize final-answer metrics and give little supervision on which evidence to keep or discard. We propose typed retrieval-enhanced learning of latent inference structure TRELLIS, a trajectory-centric multi-hop RAG framework that views retrieval as constructing a tree-structured reasoning trajectory over a typed anchor graph. TRELLIS orchestrates three modules: a schema induction agent that induces executable anchor graphs, an evidence gating & extraction agent for calibrated evidence aggregation and answer extraction, and a relative-gain rewrite agent for query reformulation guided by relative retrieval gains. All decisions and counterfactuals are logged in a causal trace ledger and optimized via phase-coupled variant of group relative policy, a group-relative policy optimization method. Experiments on multi-hop question answering and retrieval benchmarks show that TRELLIS achieves competitive or improved accuracy over strong single-hop, iterative, and multi-agent RAG baselines on several benchmarks, while yielding structured, auditable retrieval trajectories. ? 2026 IOP Publishing Ltd. All rights, including for text and data mining, AI training, and similar technologies, are reserved. This article is available under the terms of the https://publishingsupport.iopscience.iop.org/iop-standard/v1.

参考文献:

正在载入数据...

版权所有©北京联合大学 重庆维普资讯有限公司 渝B2-20050021-8 
渝公网安备 50019002500408号 违法和不良信息举报中心