登录    注册    忘记密码

详细信息

Risk-Aware Cooperative Multi-Agent Combinatorial Bandits via Conditional Value at Risk  ( EI收录)  

文献类型:期刊文献

英文题名:Risk-Aware Cooperative Multi-Agent Combinatorial Bandits via Conditional Value at Risk

作者:Li, Jiahong[1]; Ma, Nan[2]

第一作者:李佳洪

机构:[1] Beijing Union University, College of Robotics, Beijing, China; [2] Beijing University of Technology, School of Information Science and Technology, Beijing, China

第一机构:北京联合大学机器人学院

年份:2025

起止页码:7285-7290

外文期刊名:Proceedings - 2025 China Automation Congress, CAC 2025

收录:EI(收录号:20262320852102)

语种:英文

外文关键词:Autonomous agents - Decision making - Intelligent agents - Intelligent systems - Learning systems - Optimization - Risk assessment - Risk management - Sensor networks - Value engineering

摘要:In many real-world decision-making problems, multiple autonomous agents must cooperate to make sequential choices under uncertainty, while accounting for both reward and risk. In this paper, we formulate a risk-aware multi-agent combinatorial semi-bandit problem, where each agent selects actions that collectively form a joint combinatorial action. Each arm yields a random reward, and agents observe semi-bandit feedback. We adopt Conditional Value at Risk (CVaR) as the risk measure to capture worst-case outcomes. Our goal is to minimize the cumulative CVaR-regret, defined as the difference between the CVaR of the optimal joint action and that of the selected joint actions over time. To tackle this problem, we propose a novel CVaR-based Cooperative Thompson Sampling (CVaR-CoopTS) algorithm, wherein each agent maintains a Bayesian posterior for the unknown mean and variance of its local arms, and at each round samples a CVaR estimate for each base-arm. The agents then cooperate over sensor networks to select the joint action that minimizes the sum of the sampled CVaR values. We derive a finite-time upper bound on the cumulative CVaR-regret of our algorithm, which scales sublinearly in the time horizon and polynomially in the problem parameters. This establishes that our approach efficiently balances exploration and exploitation under risk. Extensive theoretical and simulation analysis demonstrate that our regret bound grows logarithmically with time, capturing the effects of the CVaR risk level and the variances of the arms. Our work provides a principled framework for risk-sensitive cooperative learning in combinatorial bandit settings. ? 2025 IEEE.

参考文献:

正在载入数据...

版权所有©北京联合大学 重庆维普资讯有限公司 渝B2-20050021-8 
渝公网安备 50019002500408号 违法和不良信息举报中心