What you’re building
A four-sentence paragraph about tides, scripted by a chat model and voiced by two speakers in one request. 49 seconds, $0.0074.
Transcript
Transcript
Host: Welcome back! Today we’re talking tides. Most people know the Moon causes them, but why do we get two high tides a day instead of just one?Guest: It comes down to ocean bulges. The Moon’s gravity pulls hardest on the ocean facing it, but it also creates a second bulge on the exact opposite side.Host: Wait, on the far side too? How does the Moon pull water away from itself?Guest: It pulls the Earth’s center harder than the far water! Because the pull is weaker out there, the ocean stretches outward, forming that second bulge.Host: Ah, so the Earth basically spins right beneath both of these bulges every day?Guest: Exactly! Coastlines rotate through both, which is why we usually see two high and two low tides roughly every twenty-five hours.
/api/v1/audio/speech endpoint accepts a list of turns as input, and each turn can name its own voice and delivery instructions. The model voices the whole conversation in one request and returns a single audio stream, so you never split the dialogue into per-line requests or stitch clips back together in order.
By the end, you will have a script that:
- Asks a chat model for a host-and-guest dialogue as structured JSON.
- Sends every turn to a Gemini TTS model in one request, with a different voice for each speaker.
- Saves the result as a WAV file you can play.
Before you start
You need:- An OpenRouter API key available as
OPENROUTER_API_KEY - Bun, or Node.js 24 or newer, both of which run TypeScript files directly
- About one cent of credits per minute of audio
podcast.ts and run it with bun podcast.ts or node podcast.ts.
This guide uses google/gemini-3.8-flash-lite-tts. Not every TTS provider supports multi-speaker input, and those that do not return a 400 rather than reading the whole dialogue in one voice. Text-to-Speech lists the supported models and the full parameter reference.
Step 1: Write the script
Ask a chat model for the dialogue and use structured outputs to get it back as a list of turns. A schema saves you from parsing speaker labels out of free text, and thedirection field gives the voice model something to act on.
Step 2: Voice the whole script in one request
Map each speaker to a voice, then send the turns asinput, with each turn’s direction as its instructions. A top-level instructions applies only to turns that omit their own, so this request leaves it out.
pcm only. Requesting mp3 returns 400 Gemini TTS only supports response_format="pcm". The 30 available voice names are listed under supported_voices in the models API.
Step 3: Save it as a WAV file
The response body is raw 16-bit PCM at 24 kHz, mono, as theContent-Type header says. Most players will not open raw PCM, so add a 44-byte WAV header in front of it:
ffmpeg -i podcast.wav podcast.mp3.
What it costs
The 49-second clip above cost $0.0074 to voice. Look up any request’s cost withGET /api/v1/generation?id=<generation ID>, using the ID from the response header. Writing the script with google/gemini-3.8-flash adds a fraction of a cent. For higher-quality speech, google/gemini-3.8-flash-tts takes the same request at a higher output price.
Troubleshooting
Check your work
The script should print a dialogue of six to eight alternating turns, then save apodcast.wav that plays two clearly different voices taking turns in the order the script lists them.