Music and Voices
Status: Approved initial tool direction; prompt format remains Candidate
Suno is the initial generation tool for both music and narration. The working pipeline is strongest when a piece is designed for two distinct voices, uses explicit spoken-word tags, and describes the emotional movement of each section.
The audio concept is developed with the storyboard rather than added after animation. Its selected structure becomes the timing and emotional guide for the Wan image-to-video shots and Palmier finish.
Voice functions
| Voice | Dramatic job |
|---|---|
| Lumi | Ground the idea, model listening, carry reassurance and wonder |
| Pip | Create momentum, ask the active question, embody surprise and play |
| Together | Deliver the hook and the shared discovery |
The functions can swap for story reasons, but the voices should never feel interchangeable.
Song brief structure
song:
episode: S01E__
letter: ""
duration_target: ""
emotional_arc: curious -> uncertain -> delighted -> warm
spoken_word:
opening: "[Lumi, softly spoken]"
pivot: "[Pip, excited spoken interjection]"
sections:
- intro
- call_and_response_verse
- chorus
- discovery_bridge
- final_chorus
- gentle_button
hook: ""
anchor_words: []
generation_record:
tool: Suno
output_mode: integrated_music_and_narration
prompt_version: ""
output_link: ""
status: Draft
For episodes where control is stronger with separate generations, preserve
distinct music_output and narration_output records and document how they are
combined in Palmier.
Rules
- Give each voice an intention and feeling, not just a name.
- Use spoken tags where story clarity matters.
- Let the chorus express the discovery, not merely repeat the letter.
- Write visual beats and song structure together.
- Preserve generation prompts and selected output links.
- Lock the chosen audio structure before final Wan shot timing and Palmier finishing.