76 voices, named for the read
Narrators, anchors, hosts, coaches and characters. The names describe delivery rather than a person — Smoky Noir Monotone, Bright Peppy Vlogger, Booming Ring Announcer — so you pick by the sound you are after.

For narration and character reads, Pika Speech is the fastest, most cost-efficient text-to-speech model on the market. Use it to turn any script into a spoken performance, read by one of over 50 voice presets. Or say it yourself by uploading a sample of your voice.
Narrators, anchors, hosts, coaches and characters. The names describe delivery rather than a person — Smoky Noir Monotone, Bright Peppy Vlogger, Booming Ring Announcer — so you pick by the sound you are after.
15,000 characters reach the model as a single request, so a full narration is one job rather than a stack of lines to render and stitch.
Full-band audio at the rate a video timeline already runs, so the read lands next to production dialogue without a resample on the way in.
A minute of speech generates in about a second. Change a word, re-run it, and compare takes in the time a slower model spends on the first one.
15,000 characters in one request. Past that the panel holds Generate and says how far over you are.
76 presets, each named for its delivery. The panel opens on Calm Documentary Narrator.
Not from this panel — it reads your script in one of the 76 presets. Cloning from a recording of your own is not part of this composer.
48 kHz audio, generated at a real-time factor of 0.02 — roughly a second of compute for a minute of speech.
No. This one speaks a script; melody, lyrics and a sung performance come from Pika Music, one tab across on the Mode row.
Yes. A preset is a fixed voice, so one choice carries the same identity across every line and every project.