Review outline / Outline
Tracing the evolution of generative TTS
An outline for a literature review connecting generations of speech models to questions about expressive control.
Published September 13, 2026
On this page
I’ve explored the historical evolution of generative models for text-to-speech. This page sets out the structure for the review; the full synthesis and annotated bibliography are not published here yet.
The question I want the review to lead toward is: what changes when a speech model needs to express emotion, rather than only produce intelligible speech?
Scope of the review
I plan to organize the literature by modeling approach, and connect each approach to the choices involved in generating expressive speech. The review will distinguish the model that represents or generates speech from the surrounding text processing, conditioning, and synthesis pipeline.
A consistent comparison
For each approach, I want to record:
- What is modeled and what the model is conditioned on.
- How speech is represented during training and generation.
- The training data and supervision required.
- How style or emotion can be specified, if supported.
- Reported evaluation, computational constraints, and limitations.
- Which claims are supported by comparable experiments.
Connecting reading to building
Fine-tuning a small TTS model has made this topic concrete for me. I want to connect the literature to a bounded follow-up experiment: keeping the speaker and text fixed while studying a clearly defined change in expressive conditioning.
The experiment design will need to specify the baseline, listening protocol, intelligibility evaluation, and how examples are selected. Audio comparisons should include transcripts and the same input text across conditions.
Publication plan
The full review will include its literature cutoff date, primary references, comparison tables, and open questions. The experiment will be a separate artifact linked to this review, with configuration, code, and results when available.
For the project context, see Small-model TTS.