OAK

Using Repeated Learner Utterances for Improving Mispronunciation Detection and Diagnosis

Metadata Downloads
Author(s)
Kim, Yoonjae
Type
Thesis
Degree
Master
Department
대학원 AI대학원
Advisor
Hong, Jin-Hyuk
Abstract
Mispronunciation detection and diagnosis (MDD) is a core component of computer-assisted pronunciation training (CAPT), but reliable MDD training usually depends on detailed pronunciation-error annotations. Such annotations are costly because non-native speech may contain substitutions, deletions, insertions, and accent-influenced realizations that require expert phonetic judgment. This thesis investigates whether repeated learner utterances collected during prompt-based CAPT practice can provide an additional weak supervision signal for MDD. The study focuses on a prompt-based CAPT setting in which Korean first-language learners of English read the same sentence multiple times while receiving feedback between attempts. Because each comparison keeps the learner and prompt fixed, earlier and later productions can be examined under a relatively stable speaker and sentence context. The proposed approach treats selected Trial 1–Trial 3 utterance pairs as noisy relative evidence: the later attempt is not assumed to be correct, but it may be more likely to contain an improved realization of the target phoneme span. A CAPT practice dataset was collected from 50 participants, each producing 15 prompts across three trials, and selected repeated-utterance pairs were used as auxiliary training data. The repeated-utterance signal is incorporated as an auxiliary preference loss on top of a connectionist temporal classification (CTC) MDD objective. On the L2-ARCTIC test-set MDD evaluation, the experiments show a modest improvement over the baseline: in the aggregate comparison, the λ = 0.01 preference-loss setting improves F1 by 1.05 percentage points and reduces phoneme error rate (PER) by 0.16 percentage points. An ordered-versus-shuffled comparison suggests that the temporal order of repeated attempts contributes to this gain. Taken together, these findings indicate that repeated CAPT practice traces can provide a complementary signal for MDD training. The comparison with a shuffled control further suggests that the observed gain is tied not only to additional learner productions, but also to their sequence across repeated attempts. Larger datasets and independent pronunciation assessment are needed to test the robustness of this signal for phone-level error modeling and to clarify its relationship to learner-level pronunciation outcomes.
URI
https://scholar.gist.ac.kr/handle/local/34546
Fulltext
http://gist.dcollection.net/common/orgView/200001023326
Alternative Author(s)
김윤재
Appears in Collections:
Dept. of AI > 3. Theses(Master)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.