OAK

Early-Phase Dynamic Penalty Parameter Selection and Convergence Acceleration of Consensus ADMM for Distributed Multi-Robot Task Assignment

Metadata Downloads
Author(s)
Jaejoon Kim
Type
Thesis
Degree
Master
Department
공과대학 기계로봇공학과
Advisor
Ahn, Hyo-Sung
Abstract
This thesis addresses Distributed Multi-Robot Task Assignment (MRTA), a core problem for efficiently allocating tasks among networked robots under limited communication. Consensus ADMM provides a systematic distributed optimization framework for MRTA, but its convergence speed is highly sensitive to the penalty parameter. To address this issue, this thesis proposes a finite-horizon Markov Decision Process (MDP)-based penalty selection method for accelerating C-ADMM-based MRTA. The proposed Learn-then-Freeze strategy uses tabular Q-learning to select the penalty pa- rameter during the early transient phase and then freezes it to maintain a fixed-parameter ADMM structure. Numerical simulations on the MURD-TAP solver show that the proposed base-freeze strat- egy reduces average iteration counts in the main tested settings when an appropriate adaptation horizon is used. The results also show that geometric-mean freezing can reduce reference-parameter dependence but may increase variability under sparse communication.
Keywords: Distributed multi-robot task assignment, Consensus ADMM, Penalty parameter selection, Reinforcement learning, Distributed optimization|본 논문은 제한된 통신 환경에서 네트워크로 연결된 로봇들에게 작업을 효율적으로 할당하기 위한
분산 다중 로봇 작업 할당(Multi-Robot Task Assignment, MRTA) 문제를 다룬다. Consensus ADMM
은 MRTA를 위한 체계적인 분산 최적화 프레임워크를 제공하지만, 수렴 속도는 패널티 파라미터 선택에
매우 민감하다. 이를 해결하기 위해 본 논문에서는 C-ADMM 기반 MRTA의 수렴을 가속하기 위한 유한
시간 마르코프 결정 과정(Markov Decision Process, MDP) 기반의 패널티 파라미터 선택 방법을 제안
한다. 제안하는 Learn-then-Freeze 전략은 초기 과도 구간에서 Tabular Q-learning을 사용하여 패널티
파라미터를 선택하고, 이후에는 이를 고정하여 고정 파라미터 ADMM 구조를 유지한다. MURD-TAP 솔
버를 대상으로 한 수치 시뮬레이션 결과, 적절한 적응 horizon이 사용될 경우 제안한 base-freeze 전략은
주요 실험 설정에서 평균 반복 횟수를 줄이는 효과를 보였다. 또한 geometric-mean freezing은 기준 파라
미터에 대한 의존성을 줄일 수 있지만, sparse communication 환경에서는 변동성을 증가시킬 수 있음을
확인하였다.
주제어: 분산 다중 로봇 작업 할당, Consensus ADMM, 패널티 파라미터 선택, 강화학습, 분산 최적화
URI
https://scholar.gist.ac.kr/handle/local/34499
Fulltext
http://gist.dcollection.net/common/orgView/200001017572
Alternative Author(s)
김재준
Appears in Collections:
Department of Mechanical and Robotics Engineering > 3. Theses(Master)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.