OAK

GIST Library Login

GIST Scholar College of Information and Computing Department of AI Convergence 1. Journal Articles

Lipreading Architecture Based on Multiple Convolutional Neural Networks for Sentence-Level Visual Speech Recognition

Metadata Downloads

Author(s): Jeon, Sanghun; Elsharkawy, Ahmed; Kim, Mun Sang

Type: Article

Citation: Sensors, v.22, no.1

Issued Date: 2022-01

Abstract: In visual speech recognition (VSR), speech is transcribed using only visual information to interpret tongue and teeth movements. Recently, deep learning has shown outstanding performance in VSR, with accuracy exceeding that of lipreaders on benchmark datasets. However, several problems still exist when using VSR systems. A major challenge is the distinction of words with similar pronunciation, called homophones; these lead to word ambiguity. Another technical limitation of traditional VSR systems is that visual information does not provide sufficient data for learning words such as “a”, “an”, “eight”, and “bin” because their lengths are shorter than 0.02 s. This report proposes a novel lipreading architecture that combines three different convolutional neural networks (CNNs; a 3D CNN, a densely connected 3D CNN, and a multi-layer feature fusion 3D CNN), which are followed by a two-layer bi-directional gated recurrent unit. The entire network was trained using connectionist temporal classification. The results of the standard automatic speech recognition evaluation metrics show that the proposed architecture reduced the character and word error rates of the baseline model by 5.681% and 11.282%, respectively, for the unseen-speaker dataset. Our proposed architecture exhibits improved performance even when visual ambiguity arises, thereby increasing VSR reliability for practical applications. © 2021 by the authors. Licensee MDPI, Basel, Switzerland.

Publisher: Multidisciplinary Digital Publishing Institute (MDPI)

ISSN: 1424-8220

DOI: 10.3390/s22010072

URI: https://scholar.gist.ac.kr/handle/local/31994

Appears in Collections:: Department of AI Convergence > 1. Journal Articles

메타데이터 간략히 보기메타데이터 전체 보기

공개 및 라이선스

공개 구분공개

qrcode

트윗하기

OAK GIST Scholar는 국립중앙도서관 OAK Repository 보급사업으로 구축되었습니다.