登录    注册    忘记密码

详细信息

Improving value function decomposition in cooperative multi-agent reinforcement learning  ( EI收录)  

文献类型:期刊文献

英文题名:Improving value function decomposition in cooperative multi-agent reinforcement learning

作者:Qin, Hao[1]; Hong, Chen[2,3]

第一作者:Qin, Hao

机构:[1] College of Smart City, Beijing Union University, Beijing, 100101, China; [2] Multi-Agent Systems Research Centre, Beijing Union University, Beijing, 100101, China; [3] College of Robotics, Beijing Union University, Beijing, 100101, China

第一机构:北京联合大学

通讯机构:[2]Multi-Agent Systems Research Centre, Beijing Union University, Beijing, 100101, China|[11417]北京联合大学;

年份:2026

卷号:703

外文期刊名:Neurocomputing

收录:EI(收录号:20263221275342);Scopus(收录号:2-s2.0-105046863996)

语种:英文

外文关键词:Computational methods - Graphic methods - Intelligent agents - Mixer circuits - Multi agent systems

摘要:Value decomposition methods have achieved strong performance in cooperative multi-agent reinforcement learning (MARL) under the centralized training with decentralized execution paradigm. However, effectively incorporating explicit, sparse, and time-varying coordination structures into monotonic value decomposition remains challenging. Although recent graph-based MARL methods have explored dynamic and group-aware interaction modeling, it is still non-trivial to construct lightweight and interpretable coordination graphs from local observations while preserving stable utility learning and decentralized execution. To address this issue, we propose the Dynamic Graph Attention Mixing Network (DGAT-MIX), a value decomposition framework for explicit dynamic coordination modeling. DGAT-MIX constructs an observation-driven coordination graph from local visibility and distance cues and performs masked multi-head graph attention over the resulting graph. The graph-enhanced representations are then injected into individual utilities through a residual coordination signal before monotonic mixing. This design introduces a sparse and interpretable relational inductive bias while retaining the original utility pathway for stable optimization. Experiments on MPE and SMACv2, together with additional analyses on selected SMAC scenarios, show that DGAT-MIX achieves competitive or superior performance compared with representative baselines, especially in heterogeneous, asymmetric, and dynamically changing cooperative tasks. ? 2026 Elsevier B.V.

参考文献:

正在载入数据...

版权所有©北京联合大学 重庆维普资讯有限公司 渝B2-20050021-8 
渝公网安备 50019002500408号 违法和不良信息举报中心