Skip to content
CodeAndBuild LogoCodeAndBuild

AI Video

Generate a Tutorial Voiceover with the ElevenLabs API

Call the ElevenLabs text-to-speech API from a Next.js route, save the audio, and keep the script editable.

CodeAndBuild Team8 min read
  • ElevenLabs
  • Voice
  • Next.js
  • API
On this page
  1. A script you can re-read
  2. Request the audio on the server
  3. Edit the take like a recording
  4. Voice, consent, and labels
  5. Stability and speed are a choice

A tutorial video fails when the voice and the cut disagree. Generate the voiceover first, then edit pictures to the words. ElevenLabs will turn a script into speech from a route on your server. The API key stays in an environment variable. The browser sends the script id, not the key, and your server returns a file you stored.

Write the script like a person talking, not like a blog introduction. Short sentences. One idea per line. Say the click before you say the reason. A 60 second tutorial is about 150 words. If you paste a 900 word article into the API you will get a nine minute narration and a video nobody finishes. Cut the script on paper first.

A script you can re-read

Keep the script in the repo next to the lesson, in a text file or a field in your CMS. You will regenerate line four without regenerating the whole course. Mark pauses with punctuation rather than with a forest of SSML until you need it. Proper names and version numbers should be written the way you want them said. "Next.js" and "v16" are easy to misread. Spell the awkward ones in the script.

lesson-01.txttext
Open the project folder. Start the dev server with npm run dev. When the page loads, click Guides. We are looking at the article layout, not the homepage. Pause on the title. That title is the headline search engines will show.

Request the audio on the server

POST to https://api.elevenlabs.io/v1/text-to-speech/{voice_id}. Send the xi-api-key header and a JSON body with the text and a model id. The response is audio, usually mpeg. Pick the voice id from your ElevenLabs account. A hardcoded id from a blog will not exist in your workspace. The model id eleven_multilingual_v2 is a common choice for English and several other languages. Check the current model list if a request says the model is unknown.

app/api/voiceover/route.tsts
export async function POST(request: Request) {
  const { text } = await request.json();
  if (typeof text !== "string" || text.length < 20 || text.length > 4000) {
    return Response.json({ error: "Script length is out of range." }, { status: 400 });
  }

  const voiceId = process.env.ELEVENLABS_VOICE_ID;
  const upstream = await fetch(
    "https://api.elevenlabs.io/v1/text-to-speech/" + voiceId,
    {
      method: "POST",
      headers: {
        "xi-api-key": process.env.ELEVENLABS_API_KEY ?? "",
        "Content-Type": "application/json",
        Accept: "audio/mpeg",
      },
      body: JSON.stringify({
        text,
        model_id: "eleven_multilingual_v2",
      }),
    },
  );

  if (!upstream.ok) {
    return Response.json({ error: "Voice request failed." }, { status: 502 });
  }

  const audio = await upstream.arrayBuffer();
  return new Response(audio, {
    headers: { "Content-Type": "audio/mpeg" },
  });
}

That route is the shape, not a public endpoint to deploy unchanged. Add authentication before it can spend your quota. Cap the text length. Do not log the script if it contains unpublished customer data, and never log the API key. Save the mp3 beside the lesson so the next page view does not call the API again.

Edit the take like a recording

  1. 01

    Listen once without pictures

    If a sentence is confusing as audio, it will not become clear with B-roll. Rewrite and regenerate that sentence.

  2. 02

    Split on paragraphs

    Generate the lesson in chunks of one paragraph. A mistake at the end should not force a new take of the opening.

  3. 03

    Match loudness

    Normalize chunks so paragraph two is not quieter than paragraph one. The listener hears the seam even when you do not.

  4. 04

    Then place pictures

    Drop the audio on the timeline first. Cut stills and clips to the nouns. Do not stretch a bad clip to fill a long sentence.

Voice, consent, and labels

Use a stock voice from your account, or a voice you have the rights to clone. Cloning a coworker because their demo sounded friendly is not a shortcut. Tell the viewer the narration is synthetic if a reasonable person would assume it is a specific human. For a faceless tutorial, a single line in the description is enough. For a channel that mixes real hosts and generated hosts, say which is which.

  • Pros: you can revise a sentence without booking a booth.
  • Cons: a regenerated paragraph can change pacing and force an edit.
  • Pros: one voice stays consistent across a course.
  • Cons: a key in client code will be copied, and the bill will follow.

Stability and speed are a choice

Voice settings change the take as much as the script does. A high stability setting sounds even and is easier to edit into a course. A lower setting sounds more performed and is harder to match when you regenerate one sentence. Pick a setting, write it next to the voice id, and keep it for the whole lesson. Speaking rate should stay close to a calm human. If you speed the file up in the editor to fit a shot, you will hear it. Shorten the sentence instead. Store the settings with the mp3 so a revision six weeks later does not arrive in a different voice. Quota is part of the design too: generating the full course on every page view will surprise you on the invoice. Generate once, commit or upload the audio, and treat the API as a studio, not as a CDN.

More guides