September 16, 2026 • 6 Min Read

How to turn a script into a natural voiceover on Melius

Melius Team
How to turn a script into a natural voiceover on Melius
Melius Team

A natural voiceover starts with the read

Start with how the line should land, because a voice that sounds human can still sound completely wrong for the work. A product explanation needs someone who sounds like they understand the product. A campaign film needs a performance. A tutorial needs a voice that doesn't turn every instruction into a sales pitch. Those are different reads, even when the words are identical.

On Melius, you turn a written script into speech on an audio node, choose the model and voice, and adjust the delivery before putting the result beside the picture. The script, the generated takes and the rest of the campaign can live on the same canvas. You're not downloading a voiceover from one tool to find out in another that the last sentence doesn't fit.

The useful distinction is between generating the performance and mixing the recording. Emotion belongs to the first. Volume in the finished video belongs to the second. Treating them as the same control is how a quiet read becomes a loud quiet read instead of a more confident one.

What to know before you generate

  1. Write for someone speaking. A sentence that reads well on a landing page can be hard to follow aloud. Shorten the clauses before trying to fix the voice.
  2. Choose the voice before refining the delivery. A naturally restrained voice and an energetic one won't respond to the same direction in the same way.
  3. Controls depend on the model. Melius carries several speech models, not one universal set of voice settings.
  4. Judge the voice against the picture. A take that sounds right alone may rush the reveal or leave the closing frame waiting.

What actually changes the sound

The voice carries the starting point

Voice selection is the first decision, not the finishing touch. Listen for the register, texture and pace that suit the piece before you spend time adjusting an unsuitable voice. If the brand needs something grounded and conversational, start there instead of trying to make an announcer sound as though they're talking across a table.

Melius includes speech models from ElevenLabs, Google, MiniMax and other providers. The audio model guide lists the available families. The point of having several is that you can try the same script against different approaches without moving the whole project.

Emotion is a performance instruction

A line can be reassuring without being slow, excited without being shouted, and serious without sounding flat. Give the model something specific to perform rather than asking it to make the result more natural.

Supported models accept expressive direction. Gemini 3.1 Flash TTS, for example, supports inline instructions and tags such as [whispering] and [laughing]. ElevenLabs models expose delivery settings that vary by model, including stability, style and pacing. These are ways to steer a take, not guarantees that every sentence will land exactly as intended.

Work on a short passage first. If the opening is supposed to sound like a recommendation from a friend, settle that read before generating the rest of the script. Otherwise you're reviewing a long recording whose first decision was already wrong.

Volume belongs in the mix

Once the read works, bring it into the multi-track editor with the video and music. Melius has clip volume controls, so you can balance the voice against the rest of the soundtrack without asking the speech model to perform again.

That separation matters. If the music covers the last word of each sentence, the problem isn't the script or the voice. Lower the music. If the speaker sounds too enthusiastic, changing the level won't fix the performance. Go back to the audio node and change the delivery.

Pitch is a separate requirement

Choosing a lower-register voice is not the same as shifting the pitch of an existing recording. Neither is changing speaking speed. If the job requires an exact pitch adjustment, treat that as a separate audio-processing requirement; the Melius controls described here don't establish a dedicated pitch-shifting feature.

It's a distinction worth making before a workflow becomes a promise to a client. Voice choice and expressive direction give you room to shape the sound, but they aren't substitutes for a specified pitch control.

Building the voiceover, step by step

  1. Put the spoken script on an audio node. Use the words you want read, with punctuation that reflects the pauses. Choose a text-to-speech model and a voice.
  2. Generate a short test. Start with the opening and the most awkward line, not the whole script. Names, abbreviations and long product descriptions are worth hearing early.
  3. Refine the delivery. Adjust the settings or expressive instructions supported by that model. Change one thing at a time so you know what improved the take.
  4. Generate the full read and place it in the cut. Check the timing against the visuals, especially reveals, transitions and the closing frame.
  5. Balance the soundtrack. Adjust clip volumes in the editor, listen to the complete sequence, and download the finished result when it works.

Why this belongs on the canvas

Melius is the Agentic OS for creative work, built around a canvas where the inputs and outputs remain visible. For voiceover work, the useful part is that the script isn't detached from the thing it describes. The product images, the video and the read belong to one piece of work, so a change to the wording doesn't need to become a hunt through exported files.

The same principle sits behind building a static ad from a style reference: make the direction explicit, keep the steps open, and correct the part that missed. For the larger campaign, see the Brief-to-Execution workflow.

Frequently asked questions

Can Melius turn written text into speech?

Yes. Audio nodes support text-to-speech models, so you can generate a spoken recording from a script and use it alongside your other creative assets.

Can I control the emotion in a voiceover?

You can steer expressive delivery using the settings and instructions supported by the selected model. The controls differ between models, and the result still needs a listen before it goes into the final cut.

Can I change the voiceover volume?

Yes. The multi-track editor has clip volume controls for balancing the generated voice against music and other audio. That changes its level in the mix, not the original performance.

Is voice selection the same as pitch control?

No. Selecting a voice changes the starting voice; pitch shifting changes the pitch of a recording. Don't treat one as evidence that the other is available.

Start with the first two sentences of a script. Get the read right on Melius, then build the rest around it.

Related Articles