Integration guides · 2026-09-27
Turkish and Spanish narration with Gemini 2.5 Pro Preview TTS
Prepare text, choose a voice and review pronunciation for Turkish and Spanish narration. Keep Gemini 2.5 Pro TTS requests separate from newer model parameters.
Fix the model version and task
Gemini 2.5 Pro Preview TTS generates audio from text. Record the full model name when preparing bilingual narration: google/gemini-2.5-pro-preview-tts on LLMTR. Do not confuse a general Gemini chat example with a speech-generation request. Use the documented TTS route rather than a chat flow expecting a text response.
Google's current general speech guide covers several model generations. Do not automatically transfer newer metadata fields or voice-design features to 2.5 Pro. This article proposes a workflow for evaluating Turkish and Spanish recordings; it does not claim live audio generation or equal measured quality in both languages. Listen to your own short trial before deciding to publish.
Finish translation before generating audio
Verify in writing that the Turkish and Spanish scripts carry the same information. Asking speech generation to translate and pronounce everything at once makes failures harder to diagnose. Check dates, numbers, people and product names first. Generate the two scripts separately so that correcting one language does not require regenerating the other.
Choose the intended audience for the Spanish script. Regional vocabulary and forms of address should be settled during editing. Clarify how abbreviations and names with Turkish suffixes should be spoken. Labels written for visual interfaces may be unclear when read aloud; use short sentences explaining what the listener should do. The scripts do not need identical word counts to convey the same message.
Verify voice selection with a short request
Use model, input and voice in LLMTR's TTS request. Gemini TTS requires voice. Start with one short sentence and confirm that the response opens as audio. The body below is a trial input, not an observed output. For the Spanish attempt, replace the input text while preserving the initial conditions.
Do not parse the response as chat JSON: this endpoint returns audio bytes. Check the file format and playback. A saved file is not proof of a finished task; an incorrect, incomplete or silent recording can also exist as a file. During the first listening pass, check that the beginning and end are intact and the complete script was read.
Short request body for POST /v1/audio/speech
{
"model": "google/gemini-2.5-pro-preview-tts",
"input": "Merhaba. Bu kısa kayıt, seslendirme kontrolü için hazırlanmıştır.",
"voice": "Kore"
}
Listen to the effect of delivery instructions
If you add a direction such as calm, clear or slower delivery, do not assume it worked merely because it appeared in the request. Listen for the intended effect. If the direction itself is spoken, its placement or interpretation may differ from what you expected. Do not invent fields outside the selected model's documented request format.
Changing voice, script, pacing request and punctuation together makes comparison difficult. Fix the text first, then change one preference. If you use a pronunciation spelling for a name, do not replace the original published text with that helper spelling. Keep speech-production text distinct from visible content where necessary, and verify that meaning and product names remain correct in the final review.
Review each language independently
Have someone who understands the language review each recording. Fluent sound does not establish accurate numbers or names. Mark dates, decimal amounts, short codes and abbreviations for attention. Do not assume that the same voice produces identical pronunciation behavior in both languages. Each language must pass its own acceptance conditions.
Record the time of each problematic segment so a correction does not require reviewing the entire script again. Associate recordings with script revision and model identifier rather than relying only on filenames. If a second attempt uses different text and sounds better, that does not demonstrate that the first text's problem was fixed.
| Area | Review question |
|---|---|
| Meaning | Was any part omitted? |
| Names | Were names pronounced consistently and correctly? |
| Numbers | Are dates and amounts understandable? |
| Rhythm | Do pauses preserve sentence meaning? |
| File | Are beginning, ending and playback correct? |
Expand only after accepting the short sample
Once the short recording passes, divide longer scripts at natural boundaries. Splitting a sentence or separating a number from its unit can damage the listening experience. Preserve voice selection and editorial conventions across segments. Listen to the combined recording too: individually good segments can still produce repetition or abrupt transitions when joined.
Do not estimate cost solely from the duration of the final file. Check the selected model's pricing units and usage for all attempts; regeneration belongs in the total. Record version and date when working with a preview model. Retest both languages after migrating. A listening approval for the previous model does not automatically approve its replacement, even if the request looks similar.
Frequently asked questions
Does Gemini TTS require a voice field?
Yes. LLMTR requires voice for Gemini TTS; a request without it is rejected with 400.
Should I translate Spanish during speech generation?
This guide recommends checking the written translation first and generating audio afterward, so translation and pronunciation errors remain distinguishable.
Was bilingual speech quality measured here?
No. This is a production and review method. Generate your own scripts and review both languages independently.