One-shots and long textures
1s for a click, an impact or a UI blip. 20s for a riser that builds the whole way, or an arena that cheers, chants and claps without a seam.

Our text-to-SFX model that creates clean, prompt-faithful sound effects, instantly usable in videos, games, and editing workflows. Use it to generate up to 20 seconds of stereo sound.
1s for a click, an impact or a UI blip. 20s for a riser that builds the whole way, or an arena that cheers, chants and claps without a seam.
4,000 characters of brief: brass gears and an escapement inside an antique clock, reverberant stadium ambience with no music, a close-mic'd pendulum a foot from your ear.
The render follows the brief without adding unrelated speech, music or background noise unless you ask for it — so an effect drops into a mix without a bed you have to strip out.
Choose one of 13 keys and it is appended to the brief, so an impact, a riser or a tonal accent lands in the same key as the track it is going under.
1 to 20 seconds, in whole seconds, on the Duration chip. Any length sends no duration at all, which leaves the model's 10-second default standing.
No. Pick a category and hit Generate — an empty description sends "sound effect" as the brief, with the category, characteristics and key still appended.
Up to 20 seconds of 44.1 kHz stereo audio, and generation averages 0.847 seconds end to end.
They are appended to your description as tags. The route takes a prompt, a duration and a seed and nothing else, so the category, the characteristics and the key are all said in the prompt.
No. The brief is followed without unrelated speech, music or background noise unless it is part of what you described.
No — this is effects and texture. Songs, arrangements and sung vocals come from Pika Music, one tab across on the Mode row.
4,000 characters, including the tags the Category, Characteristics and Key chips append. The panel says how much room is left once they are counted.