AI Video
Generate a Tutorial Voiceover with the ElevenLabs API
Call the ElevenLabs text-to-speech API from a Next.js route, save the audio, and keep the script editable.
- ElevenLabs
- Voice
- Next.js
- API
On this page
A tutorial video fails when the voice and the cut disagree. Generate the voiceover first, then edit pictures to the words. ElevenLabs will turn a script into speech from a route on your server. The API key stays in an environment variable. The browser sends the script id, not the key, and your server returns a file you stored.
Write the script like a person talking, not like a blog introduction. Short sentences. One idea per line. Say the click before you say the reason. A 60 second tutorial is about 150 words. If you paste a 900 word article into the API you will get a nine minute narration and a video nobody finishes. Cut the script on paper first.
A script you can re-read
Keep the script in the repo next to the lesson, in a text file or a field in your CMS. You will regenerate line four without regenerating the whole course. Mark pauses with punctuation rather than with a forest of SSML until you need it. Proper names and version numbers should be written the way you want them said. "Next.js" and "v16" are easy to misread. Spell the awkward ones in the script.
Open the project folder. Start the dev server with npm run dev. When the page loads, click Guides. We are looking at the article layout, not the homepage. Pause on the title. That title is the headline search engines will show.
Request the audio on the server
POST to https://api.elevenlabs.io/v1/text-to-speech/{voice_id}. Send the xi-api-key header and a JSON body with the text and a model id. The response is audio, usually mpeg. Pick the voice id from your ElevenLabs account. A hardcoded id from a blog will not exist in your workspace. The model id eleven_multilingual_v2 is a common choice for English and several other languages. Check the current model list if a request says the model is unknown.
export async function POST(request: Request) {
const { text } = await request.json();
if (typeof text !== "string" || text.length < 20 || text.length > 4000) {
return Response.json({ error: "Script length is out of range." }, { status: 400 });
}
const voiceId = process.env.ELEVENLABS_VOICE_ID;
const upstream = await fetch(
"https://api.elevenlabs.io/v1/text-to-speech/" + voiceId,
{
method: "POST",
headers: {
"xi-api-key": process.env.ELEVENLABS_API_KEY ?? "",
"Content-Type": "application/json",
Accept: "audio/mpeg",
},
body: JSON.stringify({
text,
model_id: "eleven_multilingual_v2",
}),
},
);
if (!upstream.ok) {
return Response.json({ error: "Voice request failed." }, { status: 502 });
}
const audio = await upstream.arrayBuffer();
return new Response(audio, {
headers: { "Content-Type": "audio/mpeg" },
});
}
That route is the shape, not a public endpoint to deploy unchanged. Add authentication before it can spend your quota. Cap the text length. Do not log the script if it contains unpublished customer data, and never log the API key. Save the mp3 beside the lesson so the next page view does not call the API again.
Edit the take like a recording
- 01
Listen once without pictures
If a sentence is confusing as audio, it will not become clear with B-roll. Rewrite and regenerate that sentence.
- 02
Split on paragraphs
Generate the lesson in chunks of one paragraph. A mistake at the end should not force a new take of the opening.
- 03
Match loudness
Normalize chunks so paragraph two is not quieter than paragraph one. The listener hears the seam even when you do not.
- 04
Then place pictures
Drop the audio on the timeline first. Cut stills and clips to the nouns. Do not stretch a bad clip to fill a long sentence.
Voice, consent, and labels
Use a stock voice from your account, or a voice you have the rights to clone. Cloning a coworker because their demo sounded friendly is not a shortcut. Tell the viewer the narration is synthetic if a reasonable person would assume it is a specific human. For a faceless tutorial, a single line in the description is enough. For a channel that mixes real hosts and generated hosts, say which is which.
- Pros: you can revise a sentence without booking a booth.
- Cons: a regenerated paragraph can change pacing and force an edit.
- Pros: one voice stays consistent across a course.
- Cons: a key in client code will be copied, and the bill will follow.
Stability and speed are a choice
Voice settings change the take as much as the script does. A high stability setting sounds even and is easier to edit into a course. A lower setting sounds more performed and is harder to match when you regenerate one sentence. Pick a setting, write it next to the voice id, and keep it for the whole lesson. Speaking rate should stay close to a calm human. If you speed the file up in the editor to fit a shot, you will hear it. Shorten the sentence instead. Store the settings with the mp3 so a revision six weeks later does not arrive in a different voice. Quota is part of the design too: generating the full course on every page view will surprise you on the invoice. Generate once, commit or upload the audio, and treat the API as a studio, not as a CDN.
