Cooperative Behavior by Multi-agent Reinforcement Learning with Abstractive Communication

Fast reinforcement learning approach to cooperative behavior acquisition in multi-agent system

IEEE/RSJ International Conference on Intelligent Robots and System ◽

10.1109/irds.2002.1041500 ◽

2003 ◽

Cited By ~ 2

Author(s):

Songhao Piao ◽

Bingrong Hong

Keyword(s):

Reinforcement Learning ◽

Cooperative Behavior ◽

Learning Approach ◽

Multi Agent System ◽

Agent System ◽

Multi Agent

Download Full-text

Cooperative Behavior Acquisition in Multi-agent Reinforcement Learning System Using Attention Degree

Neural Information Processing - Lecture Notes in Computer Science ◽

10.1007/978-3-642-34487-9_65 ◽

2012 ◽

pp. 537-544 ◽

Cited By ~ 3

Author(s):

Kunikazu Kobayashi ◽

Tadashi Kurano ◽

Takashi Kuremoto ◽

Masanao Obayashi

Keyword(s):

Reinforcement Learning ◽

Cooperative Behavior ◽

Learning System ◽

Multi Agent

Download Full-text

Two-stage training algorithm for AI robot soccer

PeerJ Computer Science ◽

10.7717/peerj-cs.718 ◽

2021 ◽

Vol 7 ◽

pp. e718

Author(s):

Taeyoung Kim ◽

Luiz Felipe Vecchietti ◽

Kyujin Choi ◽

Sanem Sariel ◽

Dongsoo Har

Keyword(s):

Reinforcement Learning ◽

Heterogeneous Agents ◽

Cooperative Behavior ◽

Simulation Software ◽

The Other ◽

Learning Performance ◽

Time Step ◽

Two Stage ◽

Robot Soccer ◽

Multi Agent

In multi-agent reinforcement learning, the cooperative learning behavior of agents is very important. In the field of heterogeneous multi-agent reinforcement learning, cooperative behavior among different types of agents in a group is pursued. Learning a joint-action set during centralized training is an attractive way to obtain such cooperative behavior; however, this method brings limited learning performance with heterogeneous agents. To improve the learning performance of heterogeneous agents during centralized training, two-stage heterogeneous centralized training which allows the training of multiple roles of heterogeneous agents is proposed. During training, two training processes are conducted in a series. One of the two stages is to attempt training each agent according to its role, aiming at the maximization of individual role rewards. The other is for training the agents as a whole to make them learn cooperative behaviors while attempting to maximize shared collective rewards, e.g., team rewards. Because these two training processes are conducted in a series in every time step, agents can learn how to maximize role rewards and team rewards simultaneously. The proposed method is applied to 5 versus 5 AI robot soccer for validation. The experiments are performed in a robot soccer environment using Webots robot simulation software. Simulation results show that the proposed method can train the robots of the robot soccer team effectively, achieving higher role rewards and higher team rewards as compared to other three approaches that can be used to solve problems of training cooperative multi-agent. Quantitatively, a team trained by the proposed method improves the score concede rate by 5% to 30% when compared to teams trained with the other approaches in matches against evaluation teams.

Download Full-text

Adaptation to Other Agent’s Behavior Using Meta-Strategy Learning by Collision Avoidance Simulation

Applied Sciences ◽

10.3390/app11041786 ◽

2021 ◽

Vol 11 (4) ◽

pp. 1786

Author(s):

Kensuke Miyamoto ◽

Norifumi Watanabe ◽

Yoshiyasu Takefuji

Keyword(s):

Reinforcement Learning ◽

Collision Avoidance ◽

Cooperative Behavior ◽

Behavioral Strategy ◽

Active Strategy ◽

Cooperative Tasks ◽

Behavioral Choices ◽

Multi Agent ◽

High Degree ◽

Passive Strategy

In human’s cooperative behavior, there are two strategies: a passive behavioral strategy based on others’ behaviors and an active behavioral strategy based on the objective-first. However, it is not clear how to acquire a meta-strategy to switch those strategies. The purpose of the proposed study is to create agents with the meta-strategy and to enable complex behavioral choices with a high degree of coordination. In this study, we have experimented by using multi-agent collision avoidance simulations as an example of cooperative tasks. In the experiments, we have used reinforcement learning to obtain an active strategy and a passive strategy by rewarding the interaction with agents facing each other. Furthermore, we have examined and verified the meta-strategy in situations with opponent’s strategy switched.

Download Full-text

Comparison Between Reinforcement Learning Methods with Different Goal Selections in Multi-Agent Cooperation

Journal of Advanced Computational Intelligence and Intelligent Informatics ◽

10.20965/jaciii.2017.p0917 ◽

2017 ◽

Vol 21 (5) ◽

pp. 917-929 ◽

Cited By ~ 2

Author(s):

Fumito Uwano ◽

◽

Keiki Takadama

Keyword(s):

Reinforcement Learning ◽

Learning Process ◽

Cooperative Behavior ◽

Learning Methods ◽

Q Learning ◽

Designed Experiments ◽

Multi Agent ◽

Agent Cooperation ◽

Maze Problem

This study discusses important factors for zero communication, multi-agent cooperation by comparing different modified reinforcement learning methods. The two learning methods used for comparison were assigned different goal selections for multi-agent cooperation tasks. The first method is called Profit Minimizing Reinforcement Learning (PMRL); it forces agents to learn how to reach the farthest goal, and then the agent closest to the goal is directed to the goal. The second method is called Yielding Action Reinforcement Learning (YARL); it forces agents to learn through a Q-learning process, and if the agents have a conflict, the agent that is closest to the goal learns to reach the next closest goal. To compare the two methods, we designed experiments by adjusting the following maze factors: (1) the location of the start point and goal; (2) the number of agents; and (3) the size of maze. The intensive simulations performed on the maze problem for the agent cooperation task revealed that the two methods successfully enabled the agents to exhibit cooperative behavior, even if the size of the maze and the number of agents change. The PMRL mechanism always enables the agents to learn cooperative behavior, whereas the YARL mechanism makes the agents learn cooperative behavior over a small number of learning iterations. In zero communication, multi-agent cooperation, it is important that only agents that have a conflict cooperate with each other.

Download Full-text

Multi-Agent Deep Reinforcement Learning for Decentralized Cooperative Traffic Signal Control

CICTP 2020 ◽

10.1061/9780784483053.039 ◽

2020 ◽

Author(s):

Yang Zhao ◽

Jian-Ming Hu ◽

Ming-Yang Gao ◽

Zuo Zhang

Keyword(s):

Reinforcement Learning ◽

Traffic Signal ◽

Signal Control ◽

Traffic Signal Control ◽

Multi Agent

Download Full-text

Output feedback reinforcement learning based optimal output synchronisation of heterogeneous discrete-time multi-agent systems

IET Control Theory and Applications ◽

10.1049/iet-cta.2018.6266 ◽

2019 ◽

Vol 13 (17) ◽

pp. 2866-2876

Author(s):

Syed Ali Asad Rizvi ◽

Zongli Lin

Keyword(s):

Reinforcement Learning ◽

Discrete Time ◽

Output Feedback ◽

Multi Agent Systems ◽

Agent Systems ◽

Optimal Output ◽

Multi Agent

Download Full-text

Multi-agent deep reinforcement learning with type-based hierarchical group communication

Applied Intelligence ◽

10.1007/s10489-020-02065-9 ◽

2021 ◽

Author(s):

Hao Jiang ◽

Dianxi Shi ◽

Chao Xue ◽

Yajie Wang ◽

Gongju Wang ◽

...

Keyword(s):

Reinforcement Learning ◽

Group Communication ◽

Multi Agent ◽

Hierarchical Group

Download Full-text

Multi-Agent Deep Reinforcement Learning Based Cooperative Edge Caching for Ultra-Dense Next-Generation Networks

IEEE Transactions on Communications ◽

10.1109/tcomm.2020.3044298 ◽

2020 ◽

pp. 1-1

Author(s):

Shuangwu Chen ◽

Zhen Yao ◽

Xiaofeng Jiang ◽

Jian Yang ◽

Lajos Hanzo

Keyword(s):

Reinforcement Learning ◽

Next Generation Networks ◽

Next Generation ◽

Multi Agent ◽

Edge Caching

Download Full-text

Coordinated Ramp Metering Control Based on Multi-Agent Reinforcement Learning

2020 35th Youth Academic Annual Conference of Chinese Association of Automation (YAC) ◽

10.1109/yac51587.2020.9337711 ◽

2020 ◽

Author(s):

Jiyuan Tan ◽

Qianqian Qiu ◽

Weiwei Guo

Keyword(s):

Reinforcement Learning ◽

Ramp Metering ◽

Multi Agent

Download Full-text