Macro-action-based multi-agent/robot deep reinforcement learning under partial observability

Yuchen Xiao

doi:10.17760/d20467227

ScienceGate Book Chapters

DISSERTATION

Macro-action-based multi-agent/robot deep reinforcement learning under partial observability

Yuchen Xiao

Year: 2022

DOI: 10.17760/d20467227

Get Full-Text PDF Get Analytical Report

Abstract

The state-of-the-art multi-agent reinforcement learning (MARL) methods have provided promising solutions to a variety of complex problems. Yet, these methods all assume that agents perform synchronized primitive-action executions so that they are not genuinely scalable to long-horizon real-world multi-agent/robot tasks that inherently require agents/robotsto asynchronously reason about high-level action selection at varying time durations. The Macro-Action Decentralized Partially Observable Markov Decision Process (MacDec-POMDP) is a general formalization for asynchronous decision-making under uncertainty in fully cooperative multi-agent tasks. In this thesis, we first propose a group of value-based RL approaches for MacDec-POMDPs, where agents are allowed to perform asynchronous learning and decision-making with macro-action-value functions in three paradigms: decentralized learning and control, centralized learning and control, and centralized training for decentralized execution (CTDE). Building on the above work, we formulate a set of macro-action-based policy gradient algorithms under the three training paradigms, where agents are al- lowed to directly optimize their parameterized policies in an asynchronous manner. We evaluate our methods both in simulation and on real robots over a variety of realistic domains. Empirical results demonstrate the superiority of our approaches in large multi-agent problems and validate the effectiveness of our algorithms for learning high-quality and asynchronous solutions with macro-actions.--Author's abstract

Keywords:

Reinforcement learning Computer science Observability Macro Markov decision process Artificial intelligence Variety (cybernetics) Asynchronous communication Action selection Partially observable Markov decision process Robot Machine learning Distributed computing Markov process Markov chain Agency (philosophy) Markov model Programming language

Metrics

Cited By

0.00

FWCI (Field Weighted Citation Impact)

Refs

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Reinforcement Learning in Robotics

Physical Sciences → Computer Science → Artificial Intelligence

Macro-action-based multi-agent/robot deep reinforcement learning under partial observability

Abstract

Metrics

Citation History

Topics

Related Documents

Asynchronous multi-agent deep reinforcement learning under partial observability

Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability

Coordination in Adversarial Multi-Agent with Deep Reinforcement Learning Under Partial Observability

Macro-action-based multi-agent deep reinforcement learning in cooperative tasks

Multi-Agent/Robot Deep Reinforcement Learning with Macro-Actions (Student Abstract)