| アイテムタイプ |
学術雑誌論文 / Journal Article(1) |
| 公開日 |
2012-11-07 |
| タイトル |
|
|
タイトル |
A Hidden Semi-Markov Model-Based Speech Synthesis System |
|
言語 |
en |
| 言語 |
|
|
言語 |
eng |
| 資源タイプ |
|
|
資源タイプ識別子 |
http://purl.org/coar/resource_type/c_6501 |
|
資源タイプ |
journal article |
| 著者 |
Zen, Heiga
徳田, 恵一
Masuko, Takashi
Kobayashi, Takao
Kitamura, Tadashi
|
| 著者別名 |
|
|
|
姓名 |
Tokuda, Keiichi |
|
|
言語 |
en |
|
|
姓名 |
徳田, 恵一 |
|
|
言語 |
ja |
|
|
姓名 |
トクダ, ケイイチ |
|
|
言語 |
ja-Kana |
| 著者別名 |
|
|
|
姓名 |
北村, 正 |
| 書誌情報 |
en : IEICE transactions on information and systems
巻 E90-D,
号 5,
p. 825-834,
発行日 2007-05-01
|
| 出版者 |
|
|
出版者 |
Institute of Electronics, Information and Communication Engineers |
|
言語 |
en |
| ISSN |
|
|
収録物識別子タイプ |
PISSN |
|
収録物識別子 |
0916-8532 |
| 書誌レコードID(NCID) |
|
|
収録物識別子タイプ |
NCID |
|
収録物識別子 |
AA10826272 |
| 著者版フラグ |
|
|
出版タイプ |
VoR |
|
出版タイプResource |
http://purl.org/coar/version/c_970fb48d4fbd8a85 |
| 内容記述 |
|
|
内容記述タイプ |
Other |
|
内容記述 |
A statistical speech synthesis system based on the hidden Markov model (HMM) was recently proposed. In this system, spectrum, excitation, and duration of speech are modeled simultaneously by context-dependent HMMs, and speech parameter vector sequences are generated from the HMMs themselves. This system defines a speech synthesis problem in a generative model framework and solves it based on the maximum likelihood (ML) criterion. However, there is an inconsistency: although state duration probability density functions (PDFs) are explicitly used in the synthesis part of the system, they have not been incorporated into its training part. This inconsistency can make the synthesized speech sound less natural. In this paper, we propose a statistical speech synthesis system based on a hidden semi-Markov model (HSMM), which can be viewed as an HMM with explicit state duration PDFs. The use of HSMMs can solve the above inconsistency because we can incorporate the state duration PDFs explicitly into both the synthesis and the training parts of the system. Subjective listening test results show that use of HSMMs improves the reported naturalness of synthesized speech. |
|
言語 |
en |