OAK

Sample-Efficient Learning and Compositional Generalization in Scientific Machine Learning

Metadata Downloads
Author(s)
Giup Seo
Type
Thesis
Degree
Doctor
Department
정보컴퓨팅대학 전기전자컴퓨터공학과
Advisor
Hwang, Eui Seok
Abstract
최근 과학적 머신러닝 (Scientific Machine Learning, SciML)은 복잡한 비선형 또는 고차원 편미분 방정식 (Partial differential equation, PDE)을해석하기위해머신러닝에과학적지식을결합한기법이다. 대표적인 방법으로는 물리정보기반 신경망 (Physics-Informed Neural Network, PINN)과 신경연산자 (Neural Operator, NO)가 있다. PINN에서 심층 신경망은 PDE 해를 근사하는 함수로 사용이 되는데, PDE 잔차를 손실함수에 고려함으로써 비지도학습을 통해 PDE의 해를 학습하도록 훈련된다. 이러한 학 습 방식은 데이터 의존성이 크지 않아서 학습 데이터가 부족한 상황에서 데이터의 특성이 PDE로 표현이 가능하다면 효과적으로 활용될 수 있다. 특히, PINN은 역문제나 고차원 PDE를 다룰 때 우수한 성능을 보인다. 하지만, 저차원 순방향 문제의 경우 신경망 학습 시간이 함께 고려되면 기존 수치해석 기법보다 오히려더많은연산비용이소모된다.또한,가장큰한계점으로는하나의신경망모델이일반적으로단일 초기 조건 및 경계 조건에만 적용이 가능하기 때문에 조건이 변경될 때마다 신경망 재학습이 필요하게 된다. 반면, NO는 지도학습 기반의 연산자 학습을 통해 무한차원 함수공간들 간 사상을 근사함으로써 다양한 초기 조건과 경계 조건에 대한 여러 해를 하나의 모델로 학습할 수 있다. 하지만, 초기 학습 비 용이 크다는 점과 예측할 PDE 해 데이터가 학습 데이터 분포를 벗어날 경우 성능이 급격하게 열화되는 점이 고려되어야 한다. 특히, 다양한 조건에서의 대규모 고해상도 PDE 해 데이터를 충분히 학습 과정이 필요한데, 현실적으로 연산 비용 측면에 있어서 데이터 수집에 큰 어려움이 따른다. 본 박사학위 연구는 PINN의 학습 효율성과 NO의 합성 일반화 성능을 향상시킴으로써 SciML의 주요 한계점을 극복하는 것을 목표로 한다. 먼저 PINN의 샘플 효율적인 학습을 위한 Langevin dynamics-based Adaptive Sampling (LAS)을 제안한다. PINN 학습 시, 신경망 업데이트를 위한 PDE loss 계산을 위해 collocation point가 사용되며 일반적으로주어진정의역에서 uniform random sampling방법을통해위치를결정하게된다.하지만, stiff PDEs에 적용할 경우 PDE loss landscape에서 많은 학습 정보량을 가지고 있는 high residual region이 정 의역면적대비작을때, collocation points들이주로학습정보량이적은 region들에배치가되어신경망의 수렴 속도가 느려지게 된다. 또한, 고차원 PDE의 해를 학습할 때, 메모리 비용 문제로 collocation points 들의 수를 PDE의 차원 수에 비례해서 늘리는데 어려움이 있다. 따라서, 신경망의 수렴 속도와 메모리 비용을 고려했을 때, collocation points들을 학습 정보량이 많은 곳에 위치시켜 샘플 효율적 학습을 하는 것은 매우 중요하다. LAS에서는 PDE loss를 Langevin dynamics의 잠재함수로 사용하며, collocation point의 위치를 PDE loss에 대해 gradient ascent 방향으로 이동시키고 Gaussian noise를 더한 형태로 업데이트한다. 이를 통해, LAS는 collocation points들을 학습 정보량이 높은 곳에 위치시킴으로써 Stiff PDE인 1D Allen-cahn equation에서 빠르고 안정적인 수렴을 보이고, 적은 양의 collocation points 수로 고차원 4-8D Heat equation 해를 학습할 수 있다. 다음으로 NO의 합성 일반화 개선을 위한 Spectrum-Aware Modulation (SAM) 기법을 제안한다. NO 모델을 위해 초기조건과 PDE parameter들에 대한 모든 성분을 고려하여 고해상도 PDE 학습 데이터를 생성하고 학습시키는 것은 연산 비용이 매우 크다는 한계점이 있다. 따라서, 이를 개선하기 위해 학습 데이터가 보유한 성분들로 학습 데이터가 보유하고 있지 않은 합성 데이터를 생성할 수 있는 NO 모델을 설계하는것이필요하다.본연구에서는 2D incompressible Navier-Stokes equation에대해 Fourier Neural Operator (FNO) 모델의 학습 데이터가 보유한 초기 조건 성분과 점도 성분으로 보유하고 있지 않은 합성 데이터의일반화성능을확인하고,거리가먼성분들의합성데이터생성시예측정확도가크게감소하는 것을 확인한다. 특히, 성능 저하의 원인으로 FNO의 주파수 영역에서 mode-wise transformation이 non- linear advection operator에 대해 갖는 표현력 한계에 대해서도 분석한다. 이러한 한계점들을 개선하고자, 본 연구에서 SAM은 주파수와 공간 영역의 duality property를 활용하여 linear operators들에 대해서는 PDE parameter의 spectrum의 정보를 활용하여 spectrum modulation을 적용하고, non-linear advection operators에 대해서는 spatial modulation을 적용하는 Split-Step Fourier Method (SSFM)-inspired operator splitting framework를 적용하여 FNO의 표현력을 개선한다. ©2026 서 기 업 ALL RIGHTS RESERVED|Scientific machine learning (SciML) combines machine learning and scientific knowledge to solve nonlinear or high-dimensional partial differential equations (PDEs). Physics-Informed Neural Networks (PINNs) and Neural Operators (NOs) are the representative approaches in SciML. In PINNs, Deep Neural Networks (DNNs) are trained to approximate solution functions of PDEs. PINNs are mainly trained for learning PDEs in an unsupervised manner because PDE residuals can be incorporated into the learning process. Therefore, they are effective in data-scarce scenarios and do not heavily depend on the amount of labeled data. Especially, PINNs are showing better performance in inverse problems or high-dimensional PDEs. However, PINNs face several fundamental challenges in solving forward problems for low-dimensional PDEs because training PINNs requires significantly more computational time than traditional numerical methods. In addition, the main issue is that PINNs can only be trained under a specific set of initial conditions, boundary conditions, or PDE parameters. Therefore, each model needs to be retrained from scratch when these conditions change. On the other hand, NOs can overcome these limitations by learning mappings between infinite-dimensional function spaces through operator learning. This operator learning scheme enables a single model to handle multiple initial conditions, boundary conditions, and PDE parameters. However, training NOs generally requires large-scale and high-resolution PDE datasets, and generating such datasets is computationally expensive and time-consuming. This dissertation overcomes these challenges by improving the training efficiency of PINNs and enhancing the compositional generalization of NOs. First, this dissertation proposes Langevin dynamics-based Adaptive Sampling (LAS) for sample- efficient training of PINNs. In PINN training, collocation points are used to compute PDE residual losses over a given domain, and their locations are generally determined through uniform random sampling. However, this sampling strategy becomes inefficient when informative regions occupy only a small portion of the domain, such as in stiff PDEs or PDEs with sharp transitions. In addition, in the case of high-dimensional PDEs, the number of collocation points should generally increase with dimensionality to sufficiently cover the input domain. However, this becomes impractical because of memory and computational constraints. Therefore, for both fast convergence and memory efficiency, allocating collocation points to informative regions is essential. To address this, LAS incorporates the gradient of squared PDE residuals with respect to input coordinates into the score function of Langevin dynamics. Collocation points are updated toward informative high-residual regions, while Gaussian noise is introduced to maintain stochastic exploration and prevent excessive concentration on local high-residual regions. Experimental results show that LAS achieves fast and stable convergence on benchmark PDEs, including the 1D Allen–Cahn equation. In addition, LAS demonstrates strong scalability in high-dimensional heat equations ranging from 4D to 8D using a relatively small number of collocation points. Second, this dissertation proposes Spectrum-Aware Modulation (SAM) for improving the composi- tional generalization of Neural Operators. In training NOs, collecting high-resolution training datasets for all possible combinations is impractical. Therefore, realizing compositional generalization in unseen compositions is an important research problem. To investigate whether existing NO architectures exhibit good performance in compositional generalization, Fourier Neural Operator (FNO) is evaluated on 2D incompressible Navier–Stokes equations, where unseen combinations of initial conditions and viscosities are explicitly constructed. Experimental results show that the prediction performance of vanilla FNO gradually degrades as compositions move farther away from the training distribution, especially when nonlinear advection dynamics become dominant. In addition, a fundamental limitation of spectral convolution is analyzed, where mode-wise linear transformations in the spectral domain show limited expressivity for modeling nonlinear operators because non-linear operator requires mode-mixing. SAM utilizes modulation methods to overcome these limitations. Especially, in SAM, an Split-Step Fourier Method (SSFM)-inspired operator splitting scheme effectively uses the duality between spatial and frequency domains. For this, spectral operators model linear dynamics and spatial operators compensate for nonlinear operator interactions. Experiment result shows that SAM-FNO achieves better compositional generalization across unseen combinations of initial conditions and viscosities on 2D incompressible Navier-Stokes equations. Overall, this dissertation demonstrates that incorporating physically meaningful inductive biases into both PINNs and NOs can significantly improve learning efficiency, generalization performance, and practical applicability of scientific machine learning models. ©2026 Giup Seo ALL RIGHTS RESERVED
URI
https://scholar.gist.ac.kr/handle/local/34598
Fulltext
http://gist.dcollection.net/common/orgView/200001005622
Alternative Author(s)
서기업
Appears in Collections:
Dept. of Electrical Engineering and Computer Science > 4. Theses(Ph.D)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.