OAK

Visuo-Tactile-Force-based Framework for Multi-Embodiment and Shape-Agnostic Peg-in-Hole Assembly

Metadata Downloads
Author(s)
Joosoon Lee
Type
Thesis
Degree
Doctor
Department
정보컴퓨팅대학 AI융합학과(지능로봇프로그램)
Advisor
Lee, Kyoobin
Abstract
With the rapid advancement of artificial intelligence and hardware such as humanoids and multi-fingered dexterous robot hands, the need and demand for sophisticated manipulation in diverse environments—including industrial, logistics, and domestic settings—are significantly increasing. This dissertation investigates techniques for executing peg-in-hole assembly tasks using multi-fingered dexterous robot hands. The peg-in-hole task is a fundamental and universal skill; beyond merely inserting a peg into a hole of the same shape, it can be extended to various complex operations such as assembly in manufacturing, packing in logistics, and organizing in domestic environments.
Although multi-fingered dexterous hands possess higher manipulation flexibility compared to parallel grippers, utilizing them for assembly tasks entails several challenges. Frequent physical contact between objects during the assembly process induces control uncertainty. Furthermore, severe visual occlusion caused by the close proximity of objects and the structure of the multi-fingered robot hand itself degrades vision perception performance. Therefore, this study develops a robust assembly framework that actively utilizes the 6D contact force/torque and tactile information generated during the assembly process.
Traditional modeling-based peg-in-hole methodologies are restricted to specific object shapes and require additional modeling processes for novel shapes. Moreover, existing deep learning-based visual action policy models heavily depend on the specific task environments and robot embodiments in which they were trained, requiring resource-intensive retraining for every new robot-object configuration. To overcome these scalability limitations of existing methodologies, this research proposes a unified and generalizable assembly framework across diverse object shapes and robot embodiments, along with a scalable task data collection pipeline.
To this end, this dissertation conducts the following three core studies:
(Chapter 2) Shape-Independent Peg-in-Hole Assembly Utilizing Contact Force: This chapter overcomes visual occlusion by leveraging 6D physical contact forces and develops a pose estimation AI model unconstrained by geometric shapes to execute robust assembly.
(Chapter 3) Tactile-Force-Based Multi-Robot Assembly Framework: This chapter develops a pose estimation model that fuses tactile data from the robot hand and force/torque data from the robot arm, establishing a unified assembly framework commonly applicable to various multi-fingered robot structures.
(Chapter 4) Learning Dexterous Skills from Multi-modal Human Demonstration: This chapter pursues two complementary regimes of human demonstration. The first transfers direct motion-capture–based demonstrations together with paired tactile signals into the simulator, supplying high-fidelity data for a visuo-tactile–force fusion policy. The second harvests human tasks from arbitrary monocular RGB videos, transferring the reconstructed hand–object trajectories into the simulator through mesh-based pose alignment to generate scalable robot data.
In conclusion, beyond specialized frameworks tailored to individual robot environments or specific tasks, this study establishes the foundation for a generalizable robot assembly intelligence framework capable of robustly adapting to diverse environments and multi-fingered robots based on multi-modality fusion.|휴머노이드 및 다지 로봇 핸드와 같은 하드웨어와 인공지능의 급진적인 발전에 따라 산업, 물류, 가정 등 다양한 환경에서 정교한 작업의 필요성과 수요가 크게 증가하고 있다. 본 논문에서는 다지 로봇 핸드에서 펙인홀 조립 작업을 수행하기 위한 기술을 연구한다. 펙인홀은 가장 기본적이고 범용적인 기술로, 동일한 형상의 펙을 홀에 삽입하는 작업을 넘어 제조 환경에서의 조립, 물류 환경에서의 패킹, 가정 환경에서의 정리 등 다양한 복합 작업으로 확장이 가능하다.
다지 로봇 핸드는 평행 그리퍼 대비 많은 손가락을 보유하여 높은 작업 유연성을 가지지만, 이를 활용한 조립 작업에는 여러 어려움이 따른다. 조립 과정 중 물체 간의 빈번한 물리적 접촉은 제어 불확실성을 유발하며, 물체 간의 긴밀한 접촉 및 로봇 핸드 자체의 구조로 인한 시각적 가려짐은 비전 인식 성능을 저하시킨다. 따라서 본 연구에서는 조립 과정에서 발생하는 접촉력과 촉각 정보를 적극적으로 활용하는 강건한 조립 프레임워크를 개발한다.
전통적인 모델링 기반의 펙인홀 방법론은 특정 물체의 형상에 국한되며, 새로운 형상에 대해 추가적인 모델링 과정이 요구된다. 또한 기존의 딥러닝 기반 비전 행동 정책 모델들은 학습된 특정 작업 환경 및 로봇 임바디먼트에 의존적이며, 새로운 로봇-물체 구성마다 자원 집약적인 재학습을 필요로 한다. 이러한 기존 방법론들의 확장성 한계를 극복하기 위해, 본 연구에서는 조립 물체의 형상 및 로봇 임바디먼트 전반에 일반화 가능한 조립 프레임워크와 확장성 있는 작업 데이터 수집 파이프라인을 제안한다.
이를 위해 본 논문에서는 다음의 세 가지 핵심 연구를 수행한다.
(Chapter 2) 접촉력을 활용한 형상 독립적 펙인홀 조립: 물리적 접촉력을 활용하여 시각적 가려짐을 극복하고, 기하학적 형태에 제약받지 않는 자세 추정 인공지능 모델을 개발하여 조립을 수행한다.
(Chapter 3) 촉각-힘 기반 다중 로봇 조립 프레임워크: 로봇 핸드의 촉각과 로봇 팔의 힘 데이터를 융합한 자세 추정 모델을 개발하며, 다양한 로봇 구조에 공통적으로 적용 가능한 조립 프레임워크를 구축한다.
(Chapter 4) 다중 모달리티 시연 기반 손재주 기술 학습: 모션 캡처 장비 기반의 직접적인 시연 데이터 및 비디오로 촬영된 사람의 작업을 물리 시뮬레이션 환경으로 전이하여 확장성 있는 로봇 데이터를 생성하고, 이를 기반으로 시각·촉각·힘 정보가 통합된 고도화된 조립 기술 학습 방법론을 제시한다.
결론적으로 본 연구는 개별 로봇 환경, 작업에 특화된 전문 프레임워크를 넘어, 다중 모달리티 융합을 바탕으로 다양한 환경과 다지 로봇에 강건하게 적응할 수 있는 범용적 로봇 조립 지능 프레임워크의 기반을 개발한다.
URI
https://scholar.gist.ac.kr/handle/local/34612
Fulltext
http://gist.dcollection.net/common/orgView/200001005940
Alternative Author(s)
이주순
Appears in Collections:
Dept. of AI > 4. Theses(Ph.D)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.