<?xml version="1.0" encoding="UTF-8"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns="http://purl.org/rss/1.0/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel rdf:about="https://scholar.gist.ac.kr/handle/local/7917">
    <title>Repository Collection:</title>
    <link>https://scholar.gist.ac.kr/handle/local/7917</link>
    <description />
    <items>
      <rdf:Seq>
        <rdf:li rdf:resource="https://scholar.gist.ac.kr/handle/local/34612" />
        <rdf:li rdf:resource="https://scholar.gist.ac.kr/handle/local/19852" />
        <rdf:li rdf:resource="https://scholar.gist.ac.kr/handle/local/33854" />
        <rdf:li rdf:resource="https://scholar.gist.ac.kr/handle/local/34605" />
      </rdf:Seq>
    </items>
    <dc:date>2026-09-25T17:14:09Z</dc:date>
  </channel>
  <item rdf:about="https://scholar.gist.ac.kr/handle/local/34612">
    <title>Visuo-Tactile-Force-based Framework for Multi-Embodiment and Shape-Agnostic Peg-in-Hole Assembly</title>
    <link>https://scholar.gist.ac.kr/handle/local/34612</link>
    <description>Title: Visuo-Tactile-Force-based Framework for Multi-Embodiment and Shape-Agnostic Peg-in-Hole Assembly
Author(s): Joosoon Lee
Abstract: With the rapid advancement of artificial intelligence and hardware such as humanoids and multi-fingered dexterous robot hands, the need and demand for sophisticated manipulation in diverse environments—including industrial, logistics, and domestic settings—are significantly increasing. This dissertation investigates techniques for executing peg-in-hole assembly tasks using multi-fingered dexterous robot hands. The peg-in-hole task is a fundamental and universal skill; beyond merely inserting a peg into a hole of the same shape, it can be extended to various complex operations such as assembly in manufacturing, packing in logistics, and organizing in domestic environments.
 Although multi-fingered dexterous hands possess higher manipulation flexibility compared to parallel grippers, utilizing them for assembly tasks entails several challenges. Frequent physical contact between objects during the assembly process induces control uncertainty. Furthermore, severe visual occlusion caused by the close proximity of objects and the structure of the multi-fingered robot hand itself degrades vision perception performance. Therefore, this study develops a robust assembly framework that actively utilizes the 6D contact force/torque and tactile information generated during the assembly process.
 Traditional modeling-based peg-in-hole methodologies are restricted to specific object shapes and require additional modeling processes for novel shapes. Moreover, existing deep learning-based visual action policy models heavily depend on the specific task environments and robot embodiments in which they were trained, requiring resource-intensive retraining for every new robot-object configuration. To overcome these scalability limitations of existing methodologies, this research proposes a unified and generalizable assembly framework across diverse object shapes and robot embodiments, along with a scalable task data collection pipeline.
 To this end, this dissertation conducts the following three core studies:
 (Chapter 2) Shape-Independent Peg-in-Hole Assembly Utilizing Contact Force: This chapter overcomes visual occlusion by leveraging 6D physical contact forces and develops a pose estimation AI model unconstrained by geometric shapes to execute robust assembly.
 (Chapter 3) Tactile-Force-Based Multi-Robot Assembly Framework: This chapter develops a pose estimation model that fuses tactile data from the robot hand and force/torque data from the robot arm, establishing a unified assembly framework commonly applicable to various multi-fingered robot structures.
 (Chapter 4) Learning Dexterous Skills from Multi-modal Human Demonstration: This chapter pursues two complementary regimes of human demonstration. The first transfers direct motion-capture–based demonstrations together with paired tactile signals into the simulator, supplying high-fidelity data for a visuo-tactile–force fusion policy. The second harvests human tasks from arbitrary monocular RGB videos, transferring the reconstructed hand–object trajectories into the simulator through mesh-based pose alignment to generate scalable robot data.
 In conclusion, beyond specialized frameworks tailored to individual robot environments or specific tasks, this study establishes the foundation for a generalizable robot assembly intelligence framework capable of robustly adapting to diverse environments and multi-fingered robots based on multi-modality fusion.|휴머노이드 및 다지 로봇 핸드와 같은 하드웨어와 인공지능의 급진적인 발전에 따라 산업, 물류, 가정 등 다양한 환경에서 정교한 작업의 필요성과 수요가 크게 증가하고 있다. 본 논문에서는 다지 로봇 핸드에서 펙인홀 조립 작업을 수행하기 위한 기술을 연구한다. 펙인홀은 가장 기본적이고 범용적인 기술로, 동일한 형상의 펙을 홀에 삽입하는 작업을 넘어 제조 환경에서의 조립, 물류 환경에서의 패킹, 가정 환경에서의 정리 등 다양한 복합 작업으로 확장이 가능하다.
 다지 로봇 핸드는 평행 그리퍼 대비 많은 손가락을 보유하여 높은 작업 유연성을 가지지만, 이를 활용한 조립 작업에는 여러 어려움이 따른다. 조립 과정 중 물체 간의 빈번한 물리적 접촉은 제어 불확실성을 유발하며, 물체 간의 긴밀한 접촉 및 로봇 핸드 자체의 구조로 인한 시각적 가려짐은 비전 인식 성능을 저하시킨다. 따라서 본 연구에서는 조립 과정에서 발생하는 접촉력과 촉각 정보를 적극적으로 활용하는 강건한 조립 프레임워크를 개발한다.
 전통적인 모델링 기반의 펙인홀 방법론은 특정 물체의 형상에 국한되며, 새로운 형상에 대해 추가적인 모델링 과정이 요구된다. 또한 기존의 딥러닝 기반 비전 행동 정책 모델들은 학습된 특정 작업 환경 및 로봇 임바디먼트에 의존적이며, 새로운 로봇-물체 구성마다 자원 집약적인 재학습을 필요로 한다. 이러한 기존 방법론들의 확장성 한계를 극복하기 위해, 본 연구에서는 조립 물체의 형상 및 로봇 임바디먼트 전반에 일반화 가능한 조립 프레임워크와 확장성 있는 작업 데이터 수집 파이프라인을 제안한다.
 이를 위해 본 논문에서는 다음의 세 가지 핵심 연구를 수행한다.
 (Chapter 2) 접촉력을 활용한 형상 독립적 펙인홀 조립: 물리적 접촉력을 활용하여 시각적 가려짐을 극복하고, 기하학적 형태에 제약받지 않는 자세 추정 인공지능 모델을 개발하여 조립을 수행한다.
 (Chapter 3) 촉각-힘 기반 다중 로봇 조립 프레임워크: 로봇 핸드의 촉각과 로봇 팔의 힘 데이터를 융합한 자세 추정 모델을 개발하며, 다양한 로봇 구조에 공통적으로 적용 가능한 조립 프레임워크를 구축한다.
(Chapter 4) 다중 모달리티 시연 기반 손재주 기술 학습: 모션 캡처 장비 기반의 직접적인 시연 데이터 및 비디오로 촬영된 사람의 작업을 물리 시뮬레이션 환경으로 전이하여 확장성 있는 로봇 데이터를 생성하고, 이를 기반으로 시각·촉각·힘 정보가 통합된 고도화된 조립 기술 학습 방법론을 제시한다.
 결론적으로 본 연구는 개별 로봇 환경, 작업에 특화된 전문 프레임워크를 넘어, 다중 모달리티 융합을 바탕으로 다양한 환경과 다지 로봇에 강건하게 적응할 수 있는 범용적 로봇 조립 지능 프레임워크의 기반을 개발한다.</description>
    <dc:date>2025-12-31T15:00:00Z</dc:date>
  </item>
  <item rdf:about="https://scholar.gist.ac.kr/handle/local/19852">
    <title>Training Strategies for End-to-End Noise-Robust Speech Recognition</title>
    <link>https://scholar.gist.ac.kr/handle/local/19852</link>
    <description>Title: Training Strategies for End-to-End Noise-Robust Speech Recognition
Author(s): Geon Woo Lee
Abstract: Automatic speech recognition (ASR) systems convert speech audio signals into text and are widely used in various applications. Traditional ASR consists of an acoustic model (AM) for extracting speech features and a language model (LM) for grammar and lexicon information. Recently, end-to-end (E2E) ASR models using neural networks (NN) have outperformed modular-based architectures. However, these models often perform poorly in low signal-to-noise ratio (SNR) conditions, as they are typically developed in high SNR environments. Speech enhancement (SE) or feature enhancement modules have been studied to improve low SNR performance, but they can introduce artifacts that increase error rates. Alternatively, multi-condition training (MCT) and noise-aware training (NAT) use acoustic noise as a model condition. While MCT is simple and efficient, it has limitations in low SNR conditions. Joint training of SE and ASR models has been proposed to address these issues, but conflicting gradients and frame mismatch problems make performance improvement challenging. This dissertation proposes training approaches to mitigate these joint training problems and enhance ASR performance.

First, to prevent the different tasks of the SE model and ASR model, which are two distinct models, from conflicting with each other, a training approach that separates the training procedure is proposed. The proposed training approach consists of two steps. In the first step, with the parameters of the ASR model are frozen, only the parameters of the SE model are updated using an objective function for speech quality. During this process, regularization term is applied using feature vectors extracted from the ASR encoder. Next, the parameters of both models are updated using the objective functions of SE and ASR.

Secondly, to address the conflicting gradient and frame mismatch problems, an interpreting the pipeline consisting of the SE and ASR models as a teacher-student model is proposed. In other words, the ASR model is interpreted as the teacher model to leverage linguistic knowledge, and the SE model is trained using the fine-grained those information. In addition, to transfer the frame-wise linguistic information, the acoustic tokenizer is employed as surrogate model. The acoustic tokenizer is optimized to predict cluster from k-means clustering using latent vectors of the ASR encoder. The optimized acoustic tokenizer and ASR encoder, as teacher models, transfer linguistic information to the SE model, updating the parameters of the SE model.

Finally, to mitigate problem of the cross-entropy used in the acoustic tokenizer, a pairwise distance-based loss function is proposed. In addition, to enhance the contextual representation, a contrastive learning-based relational representation between acoustic tokens and those sequence is proposed. First, samples with the same/different cluster in the acoustic tokenizer are defined as positive/negative samples, and a cluster-based pairwise distance-based loss is applied to optimize the acoustic tokenizer. Additionally, for contextual representation, contrastive learning is utilized to match the relationship between acoustic tokens and those sequence extracted from the acoustic tokenizer.

The proposed training approaches were evaluated for ASR and SE performance using simulated noisy environments and real-world audio dataset. The proposed training approaches for noise-robust ASR achieved lower word error rates (WER) in ASR performance compared to conventaional training approaches. Moreover, in SE performance, the proposed training approaches that interpreted teacher-student model achieved improved results in speech quality-related metrics compared to a separated trained SE model. These results were indicated as the addressing of the conflicting gradient and frame mismatch problems. Furthermore, comprehensive performance evaluation were conducted to verify the effectiveness of the proposed training approach for different SE and ASR models' architectures. The proposed training approaches consistently achieved better performance compared to conventional training approach in different SE model and ASR model architectures as well.</description>
    <dc:date>2023-12-31T15:00:00Z</dc:date>
  </item>
  <item rdf:about="https://scholar.gist.ac.kr/handle/local/33854">
    <title>Towards Real-World Object Recognition: From Super-Resolution to Cross-Domain and Cross-Modality</title>
    <link>https://scholar.gist.ac.kr/handle/local/33854</link>
    <description>Title: Towards Real-World Object Recognition: From Super-Resolution to Cross-Domain and Cross-Modality
Author(s): Seongmin Hwang
Abstract: Robust visual perception in real-world environments remains a fundamental challenge due to three inherent limitations of sensory data: low resolution, domain inconsistency, and modality disparity. These challenges are particularly critical in safetysensitive applications such as surveillance, autonomous driving, and defense systems, where reliable object recognition must be sustained under diverse and unpredictable conditions.
  This dissertation presents a unified research effort toward improving real-world object recognition across three complementary directions: (1) efficient super-resolution for tiny object recognition, (2) domain-generalized detection under real-world domain shifts, and (3) infrared-centric fusion for multispectral object detection. First, an efficient super-resolution framework is proposed to restore discriminative structures of small objects using kernel-attentive depthwise operations. The method achieves both computational efficiency and high reconstruction fidelity, enabling lightweight enhancement in embedded perception systems. Second, a Domain Generalized Detection Transformer (DG-DETR) is developed to enhance robustness against unseen domains. By combining wavelet-guided perturbation and domain-agnostic query selection, DG-DETR improves detection consistency across adverse weather and corruption scenarios. Finally, an Infrared-Centric Fusion (IC-Fusion) framework is introduced for multispectral detection. Through asymmetric design and cross-modal gating mechanisms, IC-Fusion effectively integrates thermal and visual cues while maintaining high efficiency.
  Extensive experiments across diverse benchmarks demonstrate that the proposed methods significantly enhance recognition accuracy and robustness under real-world conditions. Collectively, these studies contribute to advancing the reliability, generalization, and efficiency of visual perception systems, marking a step toward practical and scalable robust object recognition.|실세계 환경에서의 강인한 시각 인식은 여전히 컴퓨터 비전 분야에서 해결되지 않은 근본적인 도전 과제로 남아 있다. 이는 딥러닝 기반 인식 시스템에서 센서 데이터가 본 질적으로 가지는 세 가지 한계, 즉 저해상도, 도메인 불일치, 그리고 모달리티 간 격차 때문이다. 이러한 문제는 감시, 자율주행, 국방 시스템과 같이 안전이 중요한 응용 분 야에서 특히 치명적이며, 다양한 환경 변화 속에서도 신뢰할 수 있는 객체 인식 성능이 유지되어야 한다.
  본 논문은 이러한 문제를 해결하기 위해, 실세계 객체 인식의 강건성을 향상시키는세 가지 상호보완적 연구를 제안한다. (1) 효율적 초해상도 기반 미소 객체 인식에서는 커널 어텐션 기반 깊이분리합성곱 연산을 통해 작은 객체의 구조적 특징을 복원하면서도 연산 효율성을 확보하였다. (2) 도메인 일반화 객체 검출 연구에서는 웨이블릿 기반 특징증강 방법과 도메인 불변 쿼리 선택 기법을 결합하여, 악천후나 이미지 왜곡 등 실세계 환경 변화에도 일관된 검출 성능을 달성하였다. (3) 적외선 중심 융합 기반 다중스펙트럼 객체 검출 연구에서는 비대칭 백본 구조와 크로스 모달 게이팅 메커니즘을 통해 열화상 영상과 가시광 영상을 효율적으로 융합하면서도 높은 처리 효율을 유지하였다.
  다양한 벤치마크 실험을 통해 제안된 방법들이 실세계 환경에서의 인식 정확도와 강건성을 유의미하게 향상시킴을 확인하였다. 종합적으로 본 연구는 시각 인식 시스템의 신뢰성, 일반화 성능, 그리고 효율성을 향상시키며, 실세계 응용이 가능한 강인한 객체 인식 시스템으로 나아가기 위한 중요한 발판을 마련하였다.</description>
    <dc:date>2025-12-31T15:00:00Z</dc:date>
  </item>
  <item rdf:about="https://scholar.gist.ac.kr/handle/local/34605">
    <title>Surrogate-Assisted Evolutionary Optimization for Diffusion Models: Perspectives on Personalization and Distillation</title>
    <link>https://scholar.gist.ac.kr/handle/local/34605</link>
    <description>Title: Surrogate-Assisted Evolutionary Optimization for Diffusion Models: Perspectives on Personalization and Distillation
Author(s): Wooseok Song
Abstract: Text-to-image diffusion models have achieved strong generative quality and are widely used for image synthesis from natural language descriptions. Their practical deployment, however, requires more than high-quality generation. Pretrained models must often be adapted to individual user preferences at inference time, and large diffusion architectures must be compressed for resource-constrained environments. These two problems involve discrete deployment decisions whose evaluation requires costly procedures such as image generation, preference assessment, distillation training, or benchmark evaluation.

This dissertation studies two deployment-stage decisions in text-to-image diffusion models. The first is the ordering of a fixed prompt keyword set for inference-time personalization, and the second is the selection of recoverable distillation paths for U-Net compression. It formulates these decisions as expensive combinatorial black-box optimization problems and constructs level-specific surrogates. For the input-level search, the surrogate is built from pairwise preference evidence accumulated from evaluated prompt orders. For the structure-level search, it uses attribution-preservation features derived from cross-attention.

The personalization problem is addressed through Interactive Prompt Permutation Optimization. Model personalization is formulated as a combinatorial search problem over prompt permutations while keeping the pretrained diffusion model and the user-provided keyword set fixed. Sparse preference feedback is accumulated into an Ordering Matrix that scores candidate prompt orders through pairwise keyword precedence, and a permutation genetic algorithm searches the factorial prompt space using this surrogate. Controlled experiments show that the Ordering Matrix provides useful rank guidance and that prompt structure can steer generation toward specified preferences without model fine-tuning.

The compression problem is addressed through DELTA-Diff, a recoverability-aware distillation path optimization method. Model compression is formulated as multi-stage distillation path selection rather than parameter reduction alone. DELTA-Diff uses a hybrid surrogate that combines conventional output-level metrics with a Semantic Similarity metric derived from cross-attention attribution preservation. In Stable Diffusion U-Net compression experiments, the proposed surrogate guides the selection of intermediate architectures and improves high-ratio compression performance compared with a single-step distillation baseline.

This dissertation further develops the Cross-Attention Surrogate Design Principle. The principle states that, when a deployment objective is mediated by cross-attention-related structure, the surrogate should be constructed from the part of the model or evaluation signal most directly affected by the deployment variable. IPPO instantiates this at the input level through preference-derived pairwise ordering evidence, whereas DELTA-Diff instantiates it at the structure level through cross-attention attribution preservation. Together, the two studies indicate that level-specific surrogate design can improve the efficiency of deployment-stage optimization under limited evaluation budgets. These results provide a methodological basis for adapting pretrained diffusion models to user preferences and hardware constraints without exhaustive retraining or infeasible search.|텍스트-이미지 확산 모델은 자연어 설명으로부터 고품질 이미지를 생성하는 주요 생성 모델로 널리 사용되고 있다. 그러나 사전학습된 확산 모델을 실제 환경에 배포하기 위해서는 생성 품질의 향상만으로는 충분하지 않다. 사전학습 모델은 추론 시점에서 개별 사용자의 선호에 맞게 조정될 수 있어야 하며, 대규모 확산 모델 아키텍처는 자원 제약이 있는 환경에서 실행될 수 있도록 압축되어야 한다. 이 두 문제는 모두 이산적인 배포 의사결정 변수를 가지며, 이미지 생성, 선호 평가, 지식 증류 학습, 성능 평가와 같은 비용이 큰 절차를 통해서만 실제 목적함수를 평가할 수 있다.

본 논문은 추론 시점 개인화와 확산 모델 경량화를 배포 단계의 비용이 큰 조합적 블랙박스 최적화 문제로 정식화한다. 두 문제는 목적과 제약 조건이 서로 다르지만, 전수 탐색이 어렵고 직접 평가의 비용이 크다는 공통 구조를 가진다. 이를 해결하기 위해 본 논문은 각 배포 변수가 영향을 미치는 수준에서 사용할 수 있는 대리 신호를 구성하는 대리 모델 보조 진화 최적화 방법을 개발한다.

개인화 문제에 대해서는 대화형 프롬프트 최적화 프레임워크 IPPO를 제안한다. IPPO는 사전학습된 확산 모델과 사용자가 제공한 키워드 집합을 고정한 상태에서, 프롬프트 키워드 순서를 순열 공간 위의 조합 탐색 문제로 정식화한다. 희소한 선호 피드백은 키워드 쌍의 선후 관계를 누적하는 대리 모델로 변환되며, 순열 유전 알고리즘은 이 대리 모델을 이용해 팩토리얼 크기의 프롬프트 공간을 탐색한다. 통제된 실험 결과는 대리 모델이 유용한 순위 유도 신호를 제공하며, 모델 미세조정 없이도 프롬프트 구조를 통해 생성 결과를 사용자 표현 선호 또는 기준 기반 선호 방향으로 조정할 수 있음을 보인다.

경량화 문제에 대해서는 회복 가능성을 고려한 증류 경로 최적화 방법인 DELTA-Diff를 제안한다. DELTA-Diff는 모델 경량화를 단순한 파라미터 감소 문제가 아니라 다단계 지식 증류 경로 선택 문제로 정식화한다. 제안 방법은 일반적인 출력 수준 지표와 크로스 어텐션 맵으로부터 유도한 의미 유사도 지표를 결합한 하이브리드 대리 모델을 사용한다. 스테이블 디퓨젼 U-Net 압축 실험에서 이 대리 모델은 중간 아키텍처 선택을 안내하며, 단일 단계 지식 증류 기준선과 비교하여 높은 압축률에서의 성능을 개선한다.

본 논문은 또한 크로스 어텐션 기반 대리 신호 설계 원칙을 제안한다. 이 원칙은 배포 목적함수가 크로스 어텐션 관련 구조에 의해 매개되는 경우, 배포 변수가 주로 영향을 미치는 모델 수준 또는 평가 신호에서 대리 신호를 구성해야 한다는 것이다. IPPO는 입력 수준의 쌍별 선호 증거를 사용하고, DELTA-Diff는 구조 수준의 크로스 어텐션 귀인 보존 정보를 사용한다. 두 연구는 수준별 대리 신호 설계가 제한된 평가 예산 하에서 배포 단계 최적화의 평가 비용을 줄일 수 있음을 보인다. 이러한 결과는 사전학습된 확산 모델을 사용자 선호와 하드웨어 제약에 맞게 조정하기 위한 방법론적 기반을 제공한다.</description>
    <dc:date>2025-12-31T15:00:00Z</dc:date>
  </item>
</rdf:RDF>

