Music and Voices

Status: Approved initial tool direction; prompt format remains Candidate

Suno is the initial generation tool for both music and narration. The working pipeline is strongest when a piece is designed for two distinct voices, uses explicit spoken-word tags, and describes the emotional movement of each section.

The audio concept is developed with the storyboard rather than added after animation. Its selected structure becomes the timing and emotional guide for the Wan image-to-video shots and Palmier finish.

Voice functions

Voice Dramatic job
Lumi Ground the idea, model listening, carry reassurance and wonder
Pip Create momentum, ask the active question, embody surprise and play
Together Deliver the hook and the shared discovery

The functions can swap for story reasons, but the voices should never feel interchangeable.

Song brief structure

song:
  episode: S01E__
  letter: ""
  duration_target: ""
  emotional_arc: curious -> uncertain -> delighted -> warm
  spoken_word:
    opening: "[Lumi, softly spoken]"
    pivot: "[Pip, excited spoken interjection]"
  sections:
    - intro
    - call_and_response_verse
    - chorus
    - discovery_bridge
    - final_chorus
    - gentle_button
  hook: ""
  anchor_words: []
  generation_record:
    tool: Suno
    output_mode: integrated_music_and_narration
    prompt_version: ""
    output_link: ""
    status: Draft

For episodes where control is stronger with separate generations, preserve distinct music_output and narration_output records and document how they are combined in Palmier.

Rules

  • Give each voice an intention and feeling, not just a name.
  • Use spoken tags where story clarity matters.
  • Let the chorus express the discovery, not merely repeat the letter.
  • Write visual beats and song structure together.
  • Preserve generation prompts and selected output links.
  • Lock the chosen audio structure before final Wan shot timing and Palmier finishing.