Kyma now has a voice

Last week, Kyma started hearing. Whisper transcription and Gemini audio understanding came online behind the same single key as chat, image, and video. This week, Kyma has a voice.
Three new things you can do with the API key you already use:
- •Speak. Turn any text into hero-quality narration — 29 languages, expressive delivery, pick from 3,000+ voices.
- •Sing. Generate music from a single prompt. Describe the mood, the genre, even hum a vibe in words — get back a finished track up to five minutes long.
- •Whoosh. Generate sound effects from a description. Door slam with reverb, sword unsheathing, ambient rain at night — costs less than a stock-audio download.
All synchronous, all return audio bytes, all draw from the same balance as your other Kyma calls.
What this unlocks
The interesting part isn't the endpoints — it's what you can stitch together with one API key.
A podcast generator that writes a script with chat, narrates it in a hero voice, drops a custom intro track underneath, and punches scene transitions with sound effects. Today that requires four vendor accounts and four separate billing dashboards. On Kyma it's one balance, one ledger row per minute of output.
A voice agent that listens with Whisper, reasons with DeepSeek-V4, and replies in 75 milliseconds with the Flash voice. Real-time enough that pauses don't feel broken. One key and one balance for the whole loop — no four-vendor billing sprawl to wire up.
A video tool that generates clips with Kling or Seedance, voices the dialogue, scores the bed music, and adds foley — every step on the same key. The full audio post-production loop, prompt-driven.
That's the dog-food story in one breath: Kyma now serves chat, code, image, video, transcription, audio understanding, and audio output. Every step you'd take to make a finished piece of content can run through one balance.
The voices
Three text-to-speech models, three places they fit:
- •Multilingual v2 — the hero. Reach for this when the audio is the product. Audiobooks, brand voiceovers, narration that ships. 29 languages with consistent character across them.
- •Flash v2.5 — the conversational. Sub-100ms time-to-first-byte. Real-time agents, customer support bots, anything where pause kills the experience. Half the cost of Multilingual.
- •Turbo v2.5 — the balanced default. Faster than Multilingual, sounds better than Flash. The pick for podcast-grade narration when hero quality is overkill.
All three pull from the same 3,000-voice library. Pick a voice once, send text whenever — no fine-tuning, no enrollment.
Music: just describe it
Thirty seconds of music for the price of a coffee. Five minutes for the price of a takeout meal. Use it for podcast beds, video soundtracks, theme music for a game, or just to hear what your prompt sounds like.
Describe the mood, the genre, even hum a vibe in words — the model takes care of the rest. Lyrics support too: instrumental beds and full vocal tracks come from the same endpoint. Same for sound effects — a one-line prompt, a few seconds of audio, served back as MP3.
The full request shape and parameter list lives in the audio music API reference — no need to memorise it here.
Putting it together
You already have a Kyma key. Pick a voice from the library. Send some text. Save the file. That's it — no extra dashboard, no second billing account, no five-week procurement.
If you've been planning a project that needed audio output and stalled at "which TTS vendor do I sign up for", that decision is already made. The endpoints are live; the costs ride your existing balance; the same Authorization: Bearer ky-... header you've been using for chat now drives narration, music, and sound effects too.
What's coming next: Google's Lyria for music as a second option (cheaper at scale), Imagen 4 and Nano Banana for image, plus a few other long-tail SKUs once the licensing clears. The audio output side is live today — the next move is yours.