OAK

Learning Robust Action Policies for Robotic Manipulation under Visual Constraints

Metadata Downloads
Author(s)
Geonhyup Lee
Type
Thesis
Degree
Doctor
Department
정보컴퓨팅대학 AI융합학과
Advisor
Lee, Kyoobin
Abstract
Visual constraints in robot manipulation, including occlusion, environmental degra- dation, invisible physical quantities, and fixed viewpoint limitations, pose fundamental challenges to learning-based policies that rely primarily on visual observations. This dissertation argues that robust manipulation under such constraints requires expand- ing the robot’s access to task-relevant information rather than merely improving visual processing. To this end, it develops two complementary strategies: sensing what can- not be seen through force/torque (F/T) sensing, and seeing better by moving through active viewpoint adjustment. The dissertation presents four research contributions organized around these strate- gies. PolyFit demonstrates that force-only sensing can support robust 5-DoF peg-in- hole assembly by estimating extrinsic pose errors from multi-contact F/T observa- tions and transferring the estimator from simulation to the real world. AssemFormer extends this direction to multimodal assembly, combining vision, F/T, and proprio- ception through adaptive masking and force-guided auxiliary supervision for zero-shot sim-to-real generalization under visual degradation. ManipForce addresses the demon- stration bottleneck for contact-rich manipulation by introducing a handheld RGB–F/T data collection system and a frequency-aware transformer policy that preserves high- frequency force dynamics from human demonstrations. Finally, DynaEgo studies the complementary active-perception problem, transferring egocentric human demonstra- tions to a bimanual robot with an active head camera through simulation bridging and cross-embodiment feature alignment. Across contact-rich assembly, multimodal policy learning, human-guided manipu- lation, and active viewpoint control, the results show that policies become more robust when they are given access to the information that vision alone cannot provide. Taken together, these contributions establish force-based sensing and active perception as practical and complementary mechanisms for overcoming structural visual constraints in robot manipulation. ©2026 Geonhyup Lee ALL RIGHTS RESERVED|로봇조작에서시각정보는물체의위치,형상,주변환경을파악하기위한핵심감각 이지만, 접촉이 많은 실제 조작 상황에서는 근본적인 한계를 가진다. 작업 영역이 로봇 팔이나 말단 장치에 의해 가려지거나, 조명과 배경 조건이 변하거나, 접촉력과 마찰처 럼 영상으로는 직접 관찰할 수 없는 물리량이 조작의 성공 여부를 결정하기 때문이다. 본 논문은 이러한 시각적 제약을 단순히 더 좋은 시각 표현이나 더 많은 영상 데이터로 해결하는 데에는 한계가 있다고 보고, 로봇이 작업에 필요한 정보에 접근하는 방식을 확장하는 데 초점을 둔다. 이를위해본논문은두가지상호보완적인전략을제안한다.첫번째전략은힘/토크 (force/torque, F/T)센싱을통해시각적으로관찰할수없는접촉정보를직접감지하는 것이다. 두 번째 전략은 능동적 시점 조정을 통해 현재 시점에서는 보이지 않는 정보를 더잘관찰할수있도록카메라의위치와자세를조절하는것이다.이두전략을바탕으로 본 논문은 네 가지 연구를 제시한다. PolyFit은 다중 접촉 F/T 관측으로부터 5자유도 외재적 자세 오차를 추정하여, 시각 피드백 없이도 강건한 페그-인-홀 조립이 가능함을 보인다. AssemFormer는 시각, F/T, 고유수용감각 정보를 적응적으로 융합하여 시각 열화와 미지 형상에 강건한 zero-shot sim-to-real 조립 정책을 학습한다. ManipForce는 – iii – 핸드헬드 RGB–F/T시연수집장치와주파수인식트랜스포머정책을통해인간시연에 포함된 고주파 접촉 동역학을 로봇 조작 학습에 활용한다. DynaEgo는 인간의 자아중 심적 시연과 머리 움직임을 로봇의 능동 시점 조작으로 전이하기 위해 시뮬레이션 기반 가교와 교차 신체 특징 정렬을 사용한다. 네 연구의 결과는 로봇 조작 정책이 시각 정보만으로는 접근할 수 없는 물리적 정보 와, 고정 시점에서는 관찰하기 어려운 시각 정보를 함께 활용할 때 더 강건하고 일반화 가능한 성능을 달성할 수 있음을 보여준다. 본 논문은 힘 기반 센싱과 능동 지각이 시 각적 제약을 극복하기 위한 실용적이고 상호 보완적인 접근임을 보이며, 실제 환경에서 신뢰성 있게 작동하는 로봇 조작 시스템을 향한 하나의 방향을 제시한다. ©2026 이 건 협 ALL RIGHTS RESERVED
URI
https://scholar.gist.ac.kr/handle/local/34581
Fulltext
http://gist.dcollection.net/common/orgView/200001005927
Alternative Author(s)
이건협
Appears in Collections:
Dept. of AI > 4. Theses(Ph.D)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.