Zen Audio
Generate a full audio scene — voices, music, and sound effects — from a single prompt.
Written By Arina
Last updated 26 days ago
What it does
Zen Audio generates voice, music, and sound effects in a single pass from one prompt: describe a whole scene and get back speech, score, and sound effects already mixed together, up to 2 minutes of audio.
You can cast voices two ways:
Text-to-audio — describe the voice in words (gender, age, tone, accent).
Reference-audio voice casting — upload a short reference clip and the model matches its timbre.
Dialogue works in English and Chinese, and multi-speaker scenes give each character a distinct voice within the same generation.

Templates
A few ready-made example prompts to see the tool's range or start from something close to what you need, then edit.
Step-by-step guide
Step 1. Open Zen Audio.
Step 2. Write your prompt. Describe the scene: setting, cast, effects, voice notes, and exact lines (up to 2,048 characters). See the SCENE framework below for how to structure it.
Step 3. Add reference audio (optional). Add up to 3 clips for voice casting — each ≤ 30 seconds and ≤ 10 MB (wav/mp3/pcm/ogg_opus). Reference the clips in your prompt using @Audio1, @Audio2, @Audio3, in upload order. The clip sets the timbre — still describe emotion, tone, and pace in words.
Step 4. Use the AI Assistant (optional). Switch to the AI Assistant tab to:
Describe your scene idea and get a ready-to-use prompt drafted for you.
Ask it to improve an existing prompt — e.g. livelier voices, more atmosphere.
Ask it to build a prompt using your uploaded reference audio.
Step 5. Adjust Audio settings (optional). Expand this panel to fine-tune the output:
Format — output file format (e.g. mp3)
Sample rate — audio sample rate in Hz (e.g. 24000)
Speed — playback speed, default 1.00
Volume — output volume level, default 1.00
Pitch — pitch shift, default 0
Step 6. Generate. Click Generate (10 credits). This isn't a streaming tool — audio arrives all at once, usually after 20–60 seconds.
Step 7. Find your result under the My Generations tab.
The SCENE framework
Structure your prompt around five elements:
Setting — weather, location, acoustics: "After-school hallway, distant footsteps, reverb"
Cast — character actions, appearance: "Shouldering a backpack, waving from the door"
Effects — music mood, genre, SFX: "Deep war drums, low brass, locker 'clack'"
Voice Notes — gender, age, accent, emotion, tone, speed: "Teenage male, American accent, bright and cocky"
Exact Lines — dialogue in quotation marks: "'Hey, Emma, you free Saturday?'"
Want more templates and advanced prompt structures — timestamp control, full multi-voice/ASMR scripting, the sound tag library? Check the Guide tab inside Zen Audio for the full breakdown.
Pro tips
Write long — use most of the character limit; every detail you add informs the environment, the score, and the delivery.
Write sounds out, don't just name them — spell onomatopoeia (a zipper "zzzip", a school bell "ring-a-ling") rather than naming the sound; this is more reliable.
Match the languages — write the prompt itself in the same language as the dialogue.
Reuse reference clips for longer content — for audiobooks or episodic scenes, write scene by scene and reuse the same reference clips across generations to keep the cast consistent.
Keep reference audio CLEAR — Clean recording, Length under 30 seconds, Emotion aligned to the desired delivery, Accent consistent within each clip, Room tone steady across clips.
Every render is unique — the same prompt gives a slightly different performance each time, so it's worth generating more than once if a take isn't quite right.
Troubleshooting
Prompt too long — the limit is 2,048 characters; split long scripts into parts and generate scene by scene.
Output cut off — max output length is 2 minutes (120 seconds) per generation.
Reference audio not being cast correctly — check the clip is ≤ 30s and ≤ 10 MB, and that it's tagged in the prompt (@Audio1, @Audio2, @Audio3) in the same order it was uploaded.
Dialogue in an unsupported language — Zen Audio only speaks English and Chinese; write script lines in one of these two.
Generation feels slow — this isn't a streaming tool; audio typically arrives after 20–60 seconds.
Result doesn't match the prompt — try structuring it with the SCENE framework, be explicit about who speaks when, and spell out sound effects rather than naming them.
Results from AI models can vary and may not always be fully accurate. If you continue to experience issues, contact our support team at support@zencreator.com.
FAQ
What's next
LipSync — sync speech to a character's face on video
Pricing & Credits — see credit costs across all tools