Asynchronous multi-agent deep reinforcement learning under partial observability

Yuchen Xiao; Weihao Tan; Joshua Hoffman; Tian Xia; Christopher Amato

doi:10.1177/02783649241306124

ScienceGate Book Chapters

JOURNAL ARTICLE

Asynchronous multi-agent deep reinforcement learning under partial observability

Yuchen Xiao Weihao Tan Joshua Hoffman Tian Xia Christopher Amato

Year: 2025 Journal: The International Journal of Robotics Research Vol: 44 (8)Pages: 1257-1286 Publisher: SAGE Publishing

DOI: 10.1177/02783649241306124

Get Full-Text PDF Get Analytical Report

Abstract

The state-of-the-art multi-agent reinforcement learning (MARL) methods provide promising solutions to a variety of complex problems. Yet, these methods all assume that agents perform primitive actions in a synchronized manner, making them impractical for long-horizon real-world multi-robot tasks that inherently require robots to asynchronously reason about action selection at varying time durations. To solve this problem, we first propose a group of value-based cooperative MARL approaches for asynchronous execution using temporally extended macro-actions . Here, agents perform asynchronous learning and decision-making with macro-action-value functions in three paradigms: decentralized learning and control, centralized learning and control, and centralized training for decentralized execution (CTDE). Building on the above work, we formulate a set of macro-action-based policy gradient algorithms under the three training paradigms, where agents directly optimize their parameterized policies in an asynchronous manner. We evaluate our methods both in simulation and on real robots over a variety of realistic domains. Empirical results demonstrate the effectiveness of our algorithms for learning high-quality and asynchronous solutions with macro-actions in large multi-agent problems that were previously unsolvable via primitive-action-based approaches. The proposed approaches represent the first general MARL methods for temporally extended actions and serve as the foundation for future methods in the area.

Keywords:

Observability Reinforcement learning Asynchronous communication Computer science Artificial intelligence Reinforcement Engineering Mathematics Computer network

Metrics

Cited By

9.64

FWCI (Field Weighted Citation Impact)

Refs

0.96

Citation Normalized Percentile

Is in top 1%

Is in top 10%

Citation History

Topics

Reinforcement Learning in Robotics

Physical Sciences → Computer Science → Artificial Intelligence

Elevator Systems and Control

Physical Sciences → Engineering → Control and Systems Engineering

Adaptive Dynamic Programming Control

Physical Sciences → Computer Science → Computational Theory and Mathematics

Asynchronous multi-agent deep reinforcement learning under partial observability

Abstract

Metrics

Citation History

Topics

Related Documents

Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability

Coordination in Adversarial Multi-Agent with Deep Reinforcement Learning Under Partial Observability

Macro-action-based multi-agent/robot deep reinforcement learning under partial observability

Cooperative Multi-Agent Reinforcement Learning with Hierarchical Relation Graph under Partial Observability

Hippocampus supports multi-task reinforcement learning under partial observability.