Scaling Up Robotic Manipulation via Foundation Priors and Large-Scale Synthetic Data Sangjun Noh College of Information and Computing Gwangju Institute of Science and Technology
- Author(s)
- Sangjun Noh
- Type
- Thesis
- Degree
- Doctor
- Department
- 정보컴퓨팅대학 AI융합학과
- Advisor
- Lee, Kyoobin
- Abstract
- 로봇이 가정과 산업 현장 같은 비정형 환경에서 실질적으로 유용하게 쓰이려면, 다 양한물체를인식하고다루는조작(manipulation)능력이핵심이다.그러나기존방법은 대체로 학습 분포에 포함된 물체와 행동에 대해서만 신뢰성 있게 작동하며, 그 분포를 벗어난 미학습 물체나 더 복잡한 행동으로 일반화하는 데에는 한계가 있다. 본 학위논 문은 이러한 한계를 두 가지 과제로 정의한다. 첫째는 미학습 물체에 대한 일반화로, 실제 환경에서 로봇이 마주하는 물체는 학습 분포로 한정되지 않기 때문이다. 둘째는 행동 정책의 일반화로, 잡기와 놓기 같은 기초 동작만으로는 정밀 조립이나 양팔 다지 (dexterous)조작처럼접촉이풍부하고복잡한작업을다룰수없기때문이다.본학위논 문은이두과제를해결하기위해,웹스케일로사전학습된거대모델의지식(foundation priors)과 대규모 합성 데이터라는 두 도구를 결합하여 네 가지 방법을 제안한다. 첫번째과제에대하여,잡기와놓기는거의모든조작의출발점이지만미학습물체에 서는쉽게실패한다.기존파지검출은소규모라벨데이터에의존하고별도의물체인식 단계를 전제하여, 다양한 사용자 프롬프트에 유연하게 대응하기 어렵다. 또한 안정적 배 치는완전한 3차원모델과무게중심에대한해석적추론을요구하여,부분관측만으로는 적용이 어렵다. 이를 해결하기 위해 GraspSAM은 웹 스케일 데이터로 학습되어 물체 형상에대한일반화능력이뛰어난분할모델 SAM의지식을파지(grasping)로확장한다. 경량 모듈만을 추가로 학습하여, 점·박스·언어 등 다양한 프롬프트에 따라 미학습 물 체나 복잡한 형상까지 zero-shot으로 인식·파지하는 단일 구조를 구성한다. UOP-Net 은 물리 기반 시뮬레이션으로 구축한 대규모 합성 데이터셋으로 학습하여, 단일 시점의 부분 점군만으로 미학습 물체의 가장 안정적인 배치 평면을 예측한다. 두 방법은 미학 습 물체의 파지와 배치에서 기존 최고 수준의 성능을 달성하며, 실제 로봇 실험에서 그 일반화 성능이 검증된다. 두번째과제에대하여,정밀조립이나양팔다지조작으로나아가려면더많은시연을 모으는 것만으로는 부족하며, 행동 정책을 학습하고 일반화하는 방식 자체를 발전시켜 야 한다. 관측에서 행동으로 직접 매핑하는 방식은 로봇과 물체의 상호작용을 좌우하는 국소적 운동 단서와 미세한 접촉 동역학을 충분히 포착하지 못한다. 이를 해결하기 위해 3D Flow Diffusion Policy는 장면 수준의 3차원 흐름(3D flow)을 인지와 행동을 잇는 구조적중간표현으로삼아,흐름을먼저예측하고이에조건화된행동을추론한다.이를 통해 대규모 시뮬레이션의 다중 작업 환경에서 행동 정책 학습의 효율과 일반화를 높 인다. π-Touch는 웹 스케일로 사전학습된 시각-언어-행동(VLA) 모델 π0를 foundation prior로 삼아, 이미 강력하게 일반화된 그 시각-언어 표현 위에 촉각(힘·토크) 정보를 더함으로써, 시각과 언어만으로는 해소하기 어려운 접촉이 풍부한 정밀 조작을 다룬 다. 다만 지배적인 시각-언어 표현 위에 촉각 인코더를 처음부터 학습하면 촉각 신호가 제대로 반영되지 않으므로, 사람 시연 데이터로 촉각 인코더를 사전학습하고, 3차원 시 각-촉각 정렬(grounding)로 촉각 접촉을 시각 장면과 공간적으로 정렬한다. 이를 통해 단일 손과 양팔(bimanual) 환경에서의 다지 정밀 조작을 수행한다. 3D FDP는 대규모 시뮬레이션 다중 작업과 실제 로봇 작업에서, π-Touch는 여러 다지 조작 시뮬레이션 환경에서 기존 정책 학습 기법 대비 일관된 성능 향상을 보인다. 이상의 네 연구는 거대 사전학습 모델의 지식과 대규모 합성 데이터라는 두 도구를 두 과제에 걸쳐 적용함으로써, 비정형 환경에서 더 폭넓은 물체와 행동을 다루는 로봇 조작 연구로 나아가는 하나의 방향을 제시한다. ©2026 노 상 준 ALL RIGHTS RESERVED|For robots to be genuinely useful in unstructured environments such as households and industrial sites, they must perceive and manipulate a wide variety of objects and perform a broad range of tasks; doing so reliably, however, requires solving several gen- eralization problems. This dissertation addresses two of them as its central concerns. The first is generalization to unseen objects: real environments continually present objects that are not contained in any training set, so the robot must be able to operate on items it has never seen before. The second is generalization of the action policy: grasping and placing alone cannot accomplish the precise assembly and bimanual dex- terous behaviors that more complex, contact-rich tasks demand, so the learned policy itself must generalize beyond these simple primitives. To address these two problems, this dissertation employs two tools. The first is foundation model priors: the knowledge of segmentation and vision-language-action models trained on web-scale data is transferred into robotic manipulation, yielding broad generalization from a limited amount of robotic data. The second is large-scale synthetic data: physics-based simulation is used to obtain labels and diverse task data that are impractical to collect in the real world. Four studies spanning the two challenge axes illustrate how these tools are applied. Challenge 1: generalizing the most basic actions—grasping and placing— to unseen objects. Grasping and placing form the starting point of nearly every ma- nipulation task, yet existing methods falter on unseen objects: grasp detectors rely on small labeled datasets and assume a separate object-identification stage, which limits their flexibility under diverse prompts, while stable placement classically requires full three-dimensional object models and center-of-mass reasoning that a single partial ob- servation cannot supply. GraspSAM extends the Segment Anything Model—whose web-scale training generalizes well across object shapes—into grasping, so that unseen or geometrically complex objects can be grasped zero-shot from point, box, and lan- guage prompts, all within a single network, with its generalization confirmed on twenty unseen real-world objects. UOP-Net is trained on UOP-Sim, a large-scale synthetic dataset built through physics-based simulation, and predicts stable placement planes for unseen objects from a single-view partial point cloud, transferring to the real world without any real-world fine-tuning. Challenge 2: generalizing the action policy to diverse, contact-rich behav- iors. To move beyond grasping and placing into precise assembly and bimanual dexter- ous manipulation, it is not sufficient to collect more demonstrations; the action policy itself must generalize to these more complex behaviors. 3D Flow Diffusion Policy targets a wide range of manipulation tasks executed with a parallel-jaw (two-finger) gripper, reasoning over scene-level 3D flow as a structured intermediate representation between perception and action; this representation, learned at scale from simulated multi-task data, improves the efficiency and generalization of action learning. π-Touch builds on a π0-family vision-language-action model as a foundation prior, whose web- scale pretraining already provides strong, well-generalized vision-language features, and adds tactile (force/torque) sensing to reach the precise, contact-rich behaviors that vi- sion and language alone cannot achieve. Because a tactile encoder trained from scratch would be overwhelmed by these dominant vision-language features, π-Touch pretrains the tactile encoder on human demonstrations and spatially aligns tactile contact with the visual scene through three-dimensional visuo-tactile grounding, yielding consistent gains across the HoRA, VTDexManip, and EgoVLA simulation benchmarks in single- hand and bimanual dexterous settings. As a preliminary, simulation-only case study, real-robot validation remains future work. Taken together, these four studies apply the two tools of foundation priors and large-scale synthetic data across the two challenges—grasping and placing on unseen objects, and generalization of the action policy to diverse, contact-rich behaviors span- ning parallel-jaw multi-task manipulation to bimanual dexterous settings. This disser- tation thereby charts a direction toward robotic manipulation that handles a broader range of objects and actions in unstructured environments. ©2026 Sangjun Noh ALL RIGHTS RESERVED – iii –
- URI
- https://scholar.gist.ac.kr/handle/local/34599
- Fulltext
- http://gist.dcollection.net/common/orgView/200001005957
- 공개 및 라이선스
-
- 파일 목록
-
Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.