Music and Voices
Status: Approved initial tool direction; prompt format remains Candidate
Suno is always the first media-generation stage. After the episode concept, learning goal, and story/song brief are approved, Suno generates both music and narration. No Anima storyboard or Wan shot production begins until a specific Suno version is selected, imported into Palmier, transcribed with Whisper, and converted into a reviewed timing map.
The working pipeline is strongest when a piece is designed for two distinct voices, uses explicit spoken-word tags, and describes the emotional movement of each section.
The selected Suno audio on the Palmier timeline is the episode’s timing and emotional spine. Whisper provides the timecoded transcript used to place shot boundaries. The storyboard interprets that audio; the audio is not reshaped to justify an already-generated storyboard.
Album architecture
The 26 songs in a season must feel like one album without sounding like rewrites of one song.
Locked across the series
- The same two recognizable lead voices
- A short three-note lantern chime
- The same 6–8 second logo melody
- A brief spoken introduction
- Lumi leading melodic storytelling
- Pip leading rhythmic or playful sections
- A memorable letter refrain
- One brief guest-character vocal moment
Free to change by episode
- Melody
- Tempo
- Rhythm
- Arrangement
- Song structure
- Which voice opens or closes
- How the letter refrain returns
Locked by season
Every season receives its own:
- Instrument palette
- Rhythm vocabulary
- Genre boundaries
- Vocal-production treatment
- Emotional atmosphere
- Re-orchestrated version of the logo sting
This creates recognition at the series level, cohesion at the album level, and real musical variety at the episode level.
Voice functions
| Voice | Dramatic job |
|---|---|
| Lumi | Melodic storytelling, sustained notes, emotional bridges, wonder, observation, and occasional spoken passages |
| Pip | Rhythmic singing, quick conversational lines, rhyming responses, comic interjections, action words, and percussive letter sounds |
| Together | Alternating phrases, call-and-response, final-line harmony, and occasional unison on the letter sound |
| Guest | One short episode-specific musical identity that supports the problem without taking over the song |
The functions can swap for story reasons, but the voices should never feel interchangeable. Avoid having Lumi and Pip sing every line together. Lumi should sound warm and expressive rather than sleepy or overly delicate; Pip should sound like a real character rather than a generic children’s choir singer.
The guest moment remains brief. It may be a chant, echoed answer, rhythmic sound, musical hiccup, or whisper-sung phrase.
Suno performer-tag rule
Use voice descriptions instead of character names in performer tags. Suno may interpret a named tag as text to sing or as an unstable performer instruction.
Use stable descriptive tags:
[Spoken word — warm adult narrator]
[Young girl voice — warm, melodic]
[Young boy voice — bright, playful]
[Tiny young guest voice — determined]
[Young girl and young boy — spoken together]
Do not use:
[Lumi]
[Pip]
[Ada]
[Lumi and Pip]
Character names may appear normally inside lyrics and narration. The persistent lead descriptions should be copied verbatim between episode prompts; the guest description changes only as needed for the episode’s musical identity.
Song brief structure
song:
episode: S01E__
letter: ""
duration_target: ""
album_lock:
lead_voices: [Lumi, Pip]
lantern_chime: three_note
logo_melody_duration: 6-8_seconds
spoken_intro: true
letter_refrain: ""
season_lock:
music_bible: ""
instrument_palette: []
rhythm_vocabulary: []
genre_boundaries: []
vocal_production: ""
emotional_atmosphere: ""
logo_sting_orchestration: ""
episode_variation:
song_family: ""
tempo_feel: ""
meter: ""
arrangement: ""
structure: ""
emotional_arc: curious -> uncertain -> delighted -> warm
spoken_word:
opening: "[Young girl voice — warm, melodic; softly spoken]"
pivot: "[Young boy voice — bright, playful; excited spoken interjection]"
sections:
- intro
- call_and_response_verse
- chorus
- discovery_bridge
- final_chorus
- gentle_button
hook: ""
anchor_words: []
guest_moment:
character: ""
type: ""
text_or_sound: ""
generation_record:
tool: Suno
output_mode: integrated_music_and_narration
prompt_version: ""
output_link: ""
status: Draft
audio_lock:
selected_version: ""
duration: ""
palmier_project: ""
palmier_audio_track: ""
whisper_transcript: ""
whisper_timing_map: ""
locked_on: YYYY-MM-DD
For episodes where control is stronger with separate generations, preserve
distinct music_output and narration_output records and document how they are
combined in Palmier.
Rules
- Give each voice an intention and feeling, not just a name.
- Use descriptive performer tags, never character names, in Suno instructions.
- Use spoken-word tags where story clarity matters.
- Let the chorus express the discovery, not merely repeat the letter.
- Preserve the album locks while changing melody, tempo, rhythm, arrangement, and structure between episodes.
- Write the visual beats from the selected song and its timecoded transcript.
- Preserve generation prompts and selected output links.
- Import the chosen Suno version into Palmier before Anima storyboarding.
- Run Whisper on the exact imported song and preserve the timecoded transcript.
- Review Whisper’s words against the approved lyrics while retaining verified timing; sung words can require text correction.
- Lock the reviewed transcript and shot timing map before Anima storyboarding.
- Treat any later audio change as a dependency change that may invalidate the storyboard, Wan shots, and Palmier timeline.