OAK

From Speech Separation to Target Speaker Extraction in Deep Learning Frameworks

Metadata Downloads
Author(s)
Hyeonseung Kim
Type
Thesis
Degree
Doctor
Department
정보컴퓨팅대학 전기전자컴퓨터공학과
Advisor
Shin, Jong Won
Abstract
This dissertation investigates deep learning-based approaches for two fundamental tasks in audio processing: speech separation and target speaker extraction (TSE). Both areas aim to isolate individual speech signals from complex acoustic mixtures. The key difference is that speech separation estimates all speech signals in the input mixture, while the TSE estimates only the target speech signal using the information from the enrollment speech. In the context of speech separation with a fixed number of output channels, this dissertation proposes two novel training strategies: Choose the Best and Ignore the Rest (CBIR) and Best Matching Target (BMT). CBIR selectively computes the loss on valid output channels while disregarding others, and BMT optimally assigns target signals to outputs. Experimental results demonstrate the superiority of these strategies over conventional methods in improving separation quality when dealing with varying numbers of speakers. For the task of TSE, a novel model based on the TF-GridNet is introduced. This model leverages frequency-dependent speaker information extracted from an enrollment utterance to initialize the hidden and cell states of the network’s temporal LSTMs. Fur- thermore, cross-attention mechanisms utilizing the enrollment speech are integrated within the initial TF-GridNet blocks. The proposed TSE model achieves competitive or superior performance compared to existing methods on the WSJ0-2mix, WHAM!, and WHAMR! datasets, while maintaining lower computational complexity. High reso- lution variants of the proposed model further achieve state-of-the-art perceptual quality (PESQ scores) on these challenging datasets.|본 학위 논문에서는 딥러닝에 기반하여 음성 분리와 목표화자 추출, 두 가지 오디오 처리 기술에 대한 접근 방법을 연구한다. 두 기술 모두 복잡한 오디오 혼합 신호로부터 음성을 분리해 낸다는 공통점이 있다. 다만 음성 분리 모델은 입력 신호 내의 모든 음성 들을 각각 분리해 내는 것을 목표로 하지만, 목표화자 추출 모델은 등록 음성으로부터의 정보를 이용하여 목표 화자의 음성만을 추정한다는 차이점이 존재한다. 고정된 출력 개수를 갖는 음성 분리 모델에 대하여, 본 학위 논문에서는 최적 선택 및 나머지 무시(Choose the Best and Ignore the Rest, CBIR)와 최적 목표 할당(Best Matching Target, BMT)이라는 두 가지 훈련 방법을 제안한다. CBIR 방법은 유효한 출 력 채널에서만 손실 함수를 계산하고 다른 채널을 무시하는 방법이며, BMT 방법은 비 유효한 채널에 혼합 신호 내의 음성 신호 중 가장 최적의 신호를 할당한다. 실험 결과는 다양한 화자 수를 처리할 때 이러한 전략이 기존 방법보다 우수한 분리 품질을 달성함을 보여주었다. 목표화자 추출 방법에 대해서, TF-GridNet에 기반한 새로운 모델을 제안한다. 이 모델은 등록 화자의 음성으로부터 주파수에 따른 화자 정보를 추출하여 음성 추출 네 – iii – 트워크의 시간적 LSTM의 초기 은닉 및 셀 상태의 초깃값으로 활용한다. 또한, 등록 화자와의 교차 어텐션 메커니즘을 초기 몇 개의 블록에 적용된다. 제안된 목표화자 추출 모델은 WSJ0-2mix, WHAM!, WHAMR! 데이터셋에서 대부분의 기존 목표화자 추출 모델과 비교하여 우수하거나 유사한 성능을 보임과 동시에 더 낮은 계산 복잡도를 보 였다. 또한 모델 해상도를 증가시켰을 때, 해당 데이터셋들에서 가장 좋은 평균 PESQ 점수를 보여주었다.
URI
https://scholar.gist.ac.kr/handle/local/34572
Fulltext
http://gist.dcollection.net/common/orgView/200001005391
Alternative Author(s)
김현승
Appears in Collections:
Dept. of Electrical Engineering and Computer Science > 4. Theses(Ph.D)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.