Detailed pronunciation variant modeling for speech transcription

Author(s):  
Denis Jouvet ◽  
Dominique Fohr ◽  
Irina Illina
Keyword(s):  
Author(s):  
Haritz Arzelus ◽  
Aitor Alvarez ◽  
Conrad Bernath ◽  
Eneritz García ◽  
Emilio Granell ◽  
...  
Keyword(s):  

2016 ◽  
Author(s):  
Purushotam Radadia ◽  
Rahul Kumar ◽  
Kanika Kalra ◽  
Shirish Karande ◽  
Sachin Lodha
Keyword(s):  

2016 ◽  
Vol 81 ◽  
pp. 107-113 ◽  
Author(s):  
Rasa Lileikytė ◽  
Arseniy Gorin ◽  
Lori Lamel ◽  
Jean-Luc Gauvain ◽  
Thiago Fraga-Silva

2006 ◽  
Vol 14 (5) ◽  
pp. 1596-1608 ◽  
Author(s):  
S.F. Chen ◽  
B. Kingsbury ◽  
Lidia Mangu ◽  
D. Povey ◽  
G. Saon ◽  
...  
Keyword(s):  

2017 ◽  
Vol 68 (2) ◽  
pp. 346-354
Author(s):  
Ján Staš ◽  
Daniel Hládek ◽  
Peter Viszlay ◽  
Tomáš Koctúr

Abstract This paper describes a new Slovak speech recognition dedicated corpus built from TEDx talks and Jump Slovakia lectures. The proposed speech database consists of 220 talks and lectures in total duration of about 58 hours. Annotated speech database was generated automatically in an unsupervised manner by using acoustic speech segmentation based on principal component analysis and automatic speech transcription using two complementary speech recognition systems. The evaluation data consisting of 50 manually annotated talks and lectures in total duration of about 12 hours, has been created for evaluation of the quality of Slovak speech recognition. By unsupervised automatic annotation of TEDx talks and Jump Slovakia lectures we have obtained 21.26% of new speech segments with approximately 9.44% word error rate, suitable for retraining or adaptation of acoustic models trained beforehand.


Author(s):  
Matthias Sperber ◽  
Graham Neubig ◽  
Christian Fügen ◽  
Satoshi Nakamura ◽  
Alex Waibel
Keyword(s):  

Sign in / Sign up

Export Citation Format

Share Document