# Kyma API > Kyma API is an LLM API gateway that routes one key to 126 models and publishes measured uptime per model. One endpoint, 126 models, 70 language/reasoning, 12 retrieval, 13 image, 11 video, 20 audio. OpenAI- and Anthropic-SDK compatible, multi-route redundancy with automatic failover. Pay per token for language models, per image for image generation, per second or per clip for video, and per character / minute / track / generation for audio. Cached prompt input is billed at each model's own cached-input rate, published beside its input rate. $0.50 free credit on signup, no card required — it spends on the free tier (43 of 126 models: everything billed per token (language, embedding, reranking) at or under $3/M output and $0.50/M input, plus everything billed per unit at or under $0.01 per billing unit), not on the whole catalogue. Base URL: `https://kymaapi.com/v1` ## Docs - [Introduction](https://docs.kymaapi.com/introduction): What Kyma API is and why it exists - [Quickstart](https://docs.kymaapi.com/quickstart): Get an API key and make your first request in 30 seconds - [Pricing](https://kymaapi.com/pricing): Per-model pricing for language, image, video, and audio - [Model Recommendations](https://docs.kymaapi.com/models/recommended): Which model to use for which task - [Model Aliases](https://docs.kymaapi.com/guides/model-aliases): Use "best", "fast", "code", "cheap" instead of exact model IDs - [FAQ](https://docs.kymaapi.com/faq): Common questions and answers ## SDK and Integration Guides - [Any OpenAI Client](https://docs.kymaapi.com/guides/any-openai-client): Drop-in replacement for any OpenAI-compatible SDK or tool - [Python](https://docs.kymaapi.com/guides/python): Full Python guide with streaming, async, error handling - [JavaScript](https://docs.kymaapi.com/guides/javascript): Node.js and browser integration - [Anthropic SDK](https://docs.kymaapi.com/guides/anthropic): Use Kyma with the Anthropic Messages API - [cURL](https://docs.kymaapi.com/guides/curl): Raw HTTP examples - [LangChain](https://docs.kymaapi.com/guides/langchain): LangChain ChatOpenAI integration - [Vercel AI SDK](https://docs.kymaapi.com/guides/vercel-ai-sdk): Next.js and Vercel AI SDK provider ## AI Coding Agent Setup - [Universal Agent Setup](https://docs.kymaapi.com/guides/agent-setup): One guide for all agents. Auto-config via /v1/config endpoint - [Cline](https://kymaapi.com/for/cline): Cline (VS Code) setup with Kyma - [Roo Code](https://kymaapi.com/for/roo-code): Roo Code setup with Kyma - [Cursor](https://kymaapi.com/for/cursor): Cursor IDE setup with Kyma - [Claude Code](https://kymaapi.com/for/claude-code): Claude Code setup with Kyma - [OpenClaw](https://kymaapi.com/for/openclaw): OpenClaw setup with Kyma - [Windsurf](https://docs.kymaapi.com/guides/windsurf): Windsurf setup with Kyma - [Aider](https://docs.kymaapi.com/guides/aider): Aider CLI setup with Kyma - [LangChain](https://kymaapi.com/for/langchain): LangChain ChatOpenAI setup with Kyma - [n8n](https://kymaapi.com/for/n8n): n8n no-code automation setup with Kyma - [Open WebUI](https://kymaapi.com/for/openwebui): Open WebUI connection setup with Kyma ## Features - [Streaming](https://docs.kymaapi.com/guides/streaming): SSE streaming for all models - [Prompt Caching](https://docs.kymaapi.com/guides/prompt-caching): cached input billed at each model's own cached rate - [Tool Calling](https://docs.kymaapi.com/guides/tool-calling): Function calling across all supported models - [Structured Outputs](https://docs.kymaapi.com/guides/structured-outputs): JSON mode and response_format - [Authentication](https://docs.kymaapi.com/guides/authentication): API keys (ky-) and session tokens (ks-) - [Error Handling](https://docs.kymaapi.com/guides/error-handling): Error codes, retry logic, rate limit headers ## Site Pages - [Pricing](https://kymaapi.com/pricing): Full pricing for every model - [Models](https://kymaapi.com/models): Browse all models with live performance data - [Image Generation API](https://kymaapi.com/image-generation): FLUX, Ideogram, Recraft, Imagen - [Video Generation API](https://kymaapi.com/video-generation): Kling, Seedance, Hailuo, Veo - [Text-to-Speech API](https://kymaapi.com/text-to-speech): ElevenLabs and MiniMax voices - [Speech-to-Text API](https://kymaapi.com/transcription): Whisper transcription - [Compare](https://kymaapi.com/compare): Model-vs-model comparison pages, plus how Kyma compares to other gateways and direct APIs - [Rankings](https://kymaapi.com/rankings): Top models by tokens, speed, and uptime (live) - [Status](https://kymaapi.com/status): Live per-model availability - [Blog](https://kymaapi.com/blog): Guides on LLM gateways, coding agents, image/video/audio - [Glossary](https://kymaapi.com/glossary): LLM gateway, router, prompt caching, failover, tool calling defined ## Model Comparisons Every page below carries three things and nothing else: what each model costs on Kyma today, what Kyma measured while serving it (availability, throughput, latency), and the specification its creator publishes. No borrowed benchmark scores. Any two to four model ids from the catalogue also work as a URL, sorted and joined with `-vs-`, so a combination that is not listed here is still a page. - [grok-4.6 vs claude-fable-5](https://kymaapi.com/compare/claude-fable-5-vs-grok-4.6): frontier, side by side on price, measured uptime and published spec - [grok-4.6 vs gpt-5.6-sol](https://kymaapi.com/compare/gpt-5.6-sol-vs-grok-4.6): frontier, side by side on price, measured uptime and published spec - [claude-fable-5 vs gpt-5.6-sol-pro](https://kymaapi.com/compare/claude-fable-5-vs-gpt-5.6-sol-pro): frontier, side by side on price, measured uptime and published spec - [claude-opus-5 vs gpt-5.6-sol](https://kymaapi.com/compare/claude-opus-5-vs-gpt-5.6-sol): frontier, side by side on price, measured uptime and published spec - [grok-4.6 vs grok-4.5](https://kymaapi.com/compare/grok-4.5-vs-grok-4.6): frontier, side by side on price, measured uptime and published spec - [claude-sonnet-5 vs gemini-3.1-pro](https://kymaapi.com/compare/claude-sonnet-5-vs-gemini-3.1-pro): mid tier, side by side on price, measured uptime and published spec - [gpt-5.6-terra vs claude-sonnet-5](https://kymaapi.com/compare/claude-sonnet-5-vs-gpt-5.6-terra): mid tier, side by side on price, measured uptime and published spec - [qwen-3.8-max vs qwen-3.7-max](https://kymaapi.com/compare/qwen-3.7-max-vs-qwen-3.8-max): mid tier, side by side on price, measured uptime and published spec - [qwen-3.8-max vs deepseek-v4-pro](https://kymaapi.com/compare/deepseek-v4-pro-vs-qwen-3.8-max): mid tier, side by side on price, measured uptime and published spec - [kimi-k3 vs deepseek-v4-pro](https://kymaapi.com/compare/deepseek-v4-pro-vs-kimi-k3): mid tier, side by side on price, measured uptime and published spec - [deepseek-v4-pro vs qwen-3.6-plus](https://kymaapi.com/compare/deepseek-v4-pro-vs-qwen-3.6-plus): open-weight workhorses, side by side on price, measured uptime and published spec - [kimi-k2.6 vs minimax-m2.5](https://kymaapi.com/compare/kimi-k2.6-vs-minimax-m2.5): open-weight workhorses, side by side on price, measured uptime and published spec - [llama-3.3-70b vs gpt-oss-120b](https://kymaapi.com/compare/gpt-oss-120b-vs-llama-3.3-70b): open-weight workhorses, side by side on price, measured uptime and published spec - [qwen-3-coder vs minimax-m2.5](https://kymaapi.com/compare/minimax-m2.5-vs-qwen-3-coder): open-weight workhorses, side by side on price, measured uptime and published spec - [deepseek-r1 vs deepseek-v4-pro](https://kymaapi.com/compare/deepseek-r1-vs-deepseek-v4-pro): open-weight workhorses, side by side on price, measured uptime and published spec - [qwen-3.6-plus vs kimi-k2.6](https://kymaapi.com/compare/kimi-k2.6-vs-qwen-3.6-plus): open-weight workhorses, side by side on price, measured uptime and published spec - [glm-5.1 vs deepseek-v4-pro](https://kymaapi.com/compare/deepseek-v4-pro-vs-glm-5.1): open-weight workhorses, side by side on price, measured uptime and published spec - [qwen-3-32b vs gpt-oss-120b](https://kymaapi.com/compare/gpt-oss-120b-vs-qwen-3-32b): open-weight workhorses, side by side on price, measured uptime and published spec - [minimax-m2.5 vs glm-5.1](https://kymaapi.com/compare/glm-5.1-vs-minimax-m2.5): open-weight workhorses, side by side on price, measured uptime and published spec - [kimi-k2.6 vs deepseek-v4-pro](https://kymaapi.com/compare/deepseek-v4-pro-vs-kimi-k2.6): open-weight workhorses, side by side on price, measured uptime and published spec - [llama-3.3-70b vs qwen-3-32b](https://kymaapi.com/compare/llama-3.3-70b-vs-qwen-3-32b): open-weight workhorses, side by side on price, measured uptime and published spec - [deepseek-v4-flash vs gemini-3-flash](https://kymaapi.com/compare/deepseek-v4-flash-vs-gemini-3-flash): cheap and fast, side by side on price, measured uptime and published spec - [deepseek-v3 vs deepseek-v4-flash](https://kymaapi.com/compare/deepseek-v3-vs-deepseek-v4-flash): cheap and fast, side by side on price, measured uptime and published spec - [gemini-3-flash vs qwen-3.7-plus](https://kymaapi.com/compare/gemini-3-flash-vs-qwen-3.7-plus): cheap and fast, side by side on price, measured uptime and published spec - [qwen-3.6-plus vs deepseek-v4-flash](https://kymaapi.com/compare/deepseek-v4-flash-vs-qwen-3.6-plus): cheap and fast, side by side on price, measured uptime and published spec - [gemini-3.7-flash vs gpt-5.6-luna](https://kymaapi.com/compare/gemini-3.7-flash-vs-gpt-5.6-luna): cheap and fast, side by side on price, measured uptime and published spec - [gemini-3.7-flash vs deepseek-v4-flash](https://kymaapi.com/compare/deepseek-v4-flash-vs-gemini-3.7-flash): cheap and fast, side by side on price, measured uptime and published spec - `GET /v1/compare/popular?sort=views&window=30d`: the combinations readers open most, live, ranked by total times opened. The list above is the curated set; this endpoint is the current one. ## API Reference - [POST /v1/chat/completions](https://docs.kymaapi.com/api-reference/chat-completions): OpenAI-compatible chat completions - [GET /v1/models](https://docs.kymaapi.com/api-reference/models-list): List all models with pricing and capabilities - [POST /v1/images/generations](https://docs.kymaapi.com/api-reference/images-generations): Image generation (async) - [POST /v1/videos/generations](https://docs.kymaapi.com/api-reference/videos-generations): Video generation (async) - [POST /v1/audio/transcriptions](https://docs.kymaapi.com/api-reference/audio-transcriptions): Speech-to-text - [POST /v1/audio/speech](https://docs.kymaapi.com/api-reference/audio-speech): Text-to-speech - [POST /v1/auth/register](https://docs.kymaapi.com/api-reference/auth-register): Create account and get API key ## Discovery Endpoints (no auth required) - `GET /v1/models`: Full model list with pricing, context windows, capabilities - `GET /v1/models/recommend?usecase=coding`: Model recommendation by use case - `GET /v1/models/recommend?agent=cline`: Agent-specific recommendation with config example - `GET /v1/config`: Auto-configuration for agents (base_url, models, endpoints) - `GET /v1/capabilities`: Gateway capabilities summary - `GET /v1/credits/pricing`: Per-model pricing table - `GET /v1/status`: Service status - `GET /v1/models/uptime`: 7-day per-model uptime percentages ## Models All language models use the OpenAI chat completions format. Counts and prices below are generated from the live catalog. ### Language and reasoning models (per 1M tokens) - `claude-fable-5.1`: Claude Fable 5.1 (Anthropic), 1M context (tools, vision, reasoning). $13.50 in / $67.50 out per 1M. Anthropic's newest Mythos-class flagship — improves on Fable 5 across the board, biggest gains in agentic coding and long-running agentic workflows. 1M context, 128K output, vision, file input, adaptive thinking. - `claude-fable-5`: Claude Fable 5 (Anthropic), 1M context (tools, vision, reasoning). $13.50 in / $67.50 out per 1M. Anthropic's Mythos-class flagship tier — above Opus in capability. 1M context, vision, file input, extended reasoning. - `claude-opus-5-fast`: Claude Opus 5 Fast (Anthropic), 1M context (tools, vision, reasoning). $13.50 in / $67.50 out per 1M, cached input $1.35 per 1M. Opus 5 tuned for latency, at twice the list price. 1M context, vision, file input, prompt caching. - `claude-opus-5`: Claude Opus 5 (Anthropic), 1M context (tools, vision, reasoning). $6.75 in / $33.75 out per 1M, cached input $0.675 per 1M. Anthropic's flagship tier. 1M context, vision, file input, prompt caching, extended reasoning. - `claude-opus-4-7`: Claude Opus 4.7 (Anthropic), 1M context (tools, vision, reasoning). $6.75 in / $33.75 out per 1M, cached input $0.675 per 1M. Anthropic's flagship reasoning + agentic tier. 1M context, vision, prompt caching, code execution. Best for hardest tasks. - `kimi-k3`: Kimi K3 (Moonshot AI), 1M context (tools, vision, reasoning). $4.05 in / $20.25 out per 1M, cached input $0.405 per 1M. Moonshot's 2.8T open-weight multimodal reasoning model. Successor to K2.7 — 1M context, vision input, built for long-horizon agentic work. - `claude-sonnet-4-6`: Claude Sonnet 4.6 (Anthropic), 1M context (tools, vision, reasoning). $4.05 in / $20.25 out per 1M, cached input $0.405 per 1M. Anthropic's balanced tier. 1M context, vision, agentic tool use, prompt caching. The default workhorse. - `sonar-pro`: Sonar Pro (Perplexity), 200K context (vision). $4.05 in / $20.25 out per 1M, prompt caching supported (cached rate not published). Perplexity's pro web-search model. Deeper multi-step search, 200K context, longer cited answers. Per-request search fee on top of tokens. - `grok-4.5`: Grok 4.5 (xAI), 500K context (tools, vision, reasoning). $2.73 in / $8.189 out per 1M. xAI's current flagship. 500K context, vision, file input, reasoning; the rate doubles above 200K prompt tokens. - `grok-4.6`: Grok 4.6 (xAI), 500K context (tools, vision, reasoning). $2.70 in / $8.10 out per 1M. xAI's newest flagship. 500K context, vision, file input, reasoning; the rate doubles above 200K prompt tokens. - `gpt-5.6-sol-pro`: GPT-5.6 Sol Pro (OpenAI), 1M context (tools, vision, reasoning). $2.70 in / $13.50 out per 1M, cached input $0.3375 per 1M. Sol at the same list price with the Pro serving profile. 1.05M context, vision, file input; the rate doubles above 272K prompt tokens. - `gpt-5.6-sol`: GPT-5.6 Sol (OpenAI), 1M context (tools, vision, reasoning). $2.70 in / $13.50 out per 1M, cached input $0.3375 per 1M. OpenAI's top 5.6 tier. 1.05M context, vision, file input; the rate doubles above 272K prompt tokens. - `gpt-5.6-terra-pro`: GPT-5.6 Terra Pro (OpenAI), 1M context (tools, vision, reasoning). $2.70 in / $16.20 out per 1M, cached input $0.27 per 1M. Terra's Pro tier, twice Terra's list price. 1.05M context, vision, file input; the rate doubles above 272K prompt tokens. - `claude-sonnet-5`: Claude Sonnet 5 (Anthropic), 1M context (tools, vision, reasoning). $2.70 in / $13.50 out per 1M, cached input $0.27 per 1M. Anthropic's balanced tier. 1M context, vision, file input, prompt caching, extended reasoning. - `gpt-5.6-terra`: GPT-5.6 Terra (OpenAI), 1M context (tools, vision, reasoning). $2.70 in / $16.20 out per 1M, cached input $0.27 per 1M. OpenAI's balanced GPT-5.6 tier, between the Sol flagship and the Luna cost tier. 1M context, accepts text, images and files. - `gemini-3.1-pro`: Gemini 3.1 Pro (Google), 1M context (tools, vision, reasoning). $2.70 in / $16.20 out per 1M, cached input $0.27 per 1M. Google's Pro reasoning tier — the step above Flash. 1M context, 64K max output, accepts text, image, audio, video and file; the rate steps up above 200K prompt tokens. - `qwen-3.7-max`: Qwen 3.7 Max (Alibaba), 1M context (tools, reasoning). $2.304 in / $6.909 out per 1M, cached input $0.16875 per 1M. Alibaba's newest closed-weight flagship. 1M context, top reasoning + multilingual. - `qwen-3.8-max`: Qwen 3.8 Max (Alibaba), 1M context (tools, vision, reasoning). $2.2275 in / $6.684 out per 1M, cached input $0.2781 per 1M. Alibaba's newest flagship, GA successor to the 3.7 Max line — and the first Max with vision. 1M context, multimodal reasoning over text and images, cheaper per token than the 3.7 it replaces. - `gemini-3.5-flash`: Gemini 3.5 Flash (Google), 1M context (tools, vision, reasoning). $2.025 in / $12.15 out per 1M, cached input $0.2025 per 1M. Newest Gemini Flash. 1M context, multimodal input. - `grok-4.3`: Grok 4.3 (xAI), 1M context (tools, vision, reasoning). $1.92 in / $3.838 out per 1M, cached input $0.27 per 1M. xAI's frontier model. 1M context, strong reasoning + tool use. - `glm-5.3`: GLM 5.3 (Zhipu AI), 1M context (tools, reasoning). $1.89 in / $5.94 out per 1M, prompt caching supported (cached rate not published). Zhipu's newest flagship — coding/agentic upgrade over GLM 5.2. ~1.3M context, open weights. - `glm-5.1`: GLM 5.1 (Zhipu AI), 203K context (tools, reasoning). $1.89 in / $5.94 out per 1M, cached input $0.27675 per 1M. #1 SWE-Bench Pro open-weight. 8-hour agentic runs. - `grok-4.20-multi-agent`: Grok 4.20 Multi-Agent (xAI), 2M context (vision, reasoning). $1.767 in / $3.534 out per 1M. Grok 4.20 run as a multi-agent ensemble, at the same list price. 2M context, vision, file input, reasoning. - `grok-4.20`: Grok 4.20 (xAI), 2M context (tools, vision, reasoning). $1.763 in / $3.528 out per 1M. xAI's 4.20 generation. 2M context, vision, file input, reasoning; the rate doubles above 200K prompt tokens. - `muse-spark-1.2`: Muse Spark 1.2 (Meta), 1M context (tools, vision, reasoning). $1.688 in / $5.738 out per 1M. Meta's newest Muse Spark — the model behind Muse Code, with bigger gains in coding and agentic work over 1.1. 1M context, multimodal input (text, image, video, file, audio), reasoning. - `muse-spark-1.1`: Muse Spark 1.1 (Meta), 1M context (tools, vision, reasoning). $1.542 in / $5.243 out per 1M. Meta's Muse Spark tier. 1M context, vision, file input, reasoning — it spends ~160 tokens thinking before a one-word answer. - `grok-build`: Grok Build (xAI), 256K context (tools, vision, reasoning). $1.389 in / $2.777 out per 1M, cached input $0.27 per 1M. xAI's coding-specialized model. Fast, tool-native, built for agentic dev. - `claude-haiku-4-5`: Claude Haiku 4.5 (Anthropic), 200K context (tools, vision, reasoning). $1.35 in / $6.75 out per 1M, cached input $0.135 per 1M. Anthropic's fast tier. Sub-second TTFT, vision, tool use, prompt caching. 200K context. - `sonar`: Sonar (Perplexity), 127K context (vision). $1.35 in / $1.35 out per 1M, prompt caching supported (cached rate not published). Perplexity's live web-search model. Returns current, cited answers — grounded in a real-time search of the web. Bills a small per-request search fee on top of tokens. - `hermes-3-405b`: Hermes 3 405B (Nousresearch), 131K context (tools). $1.35 in / $1.35 out per 1M. Nous Research's fine-tune of Llama 3.1 405B, tuned for steerability rather than refusal: it follows a system prompt further than the Instruct model it is built on, which is why roleplay, persona and agent-scaffold work keep reaching for it. Reliable function calling and structured output are part of the tune, not bolted on. - `glm-5.2`: GLM 5.2 (Zhipu AI), 1M context (tools, reasoning). $1.101 in / $3.76 out per 1M, cached input $0.29835 per 1M. Frontier open-weight. #1 Intelligence Index among open models. 1M context, long-horizon agentic coding. - `gemini-3.8-flash`: Gemini 3.8 Flash (Google), 1M context (tools, vision, reasoning). $1.013 in / $5.063 out per 1M, prompt caching supported (cached rate not published). Google's newest Flash tier, GA 2026-09-02 — significant gains over 3.7 Flash on software engineering and agentic tasks. 1M context, 64K output, text, image, audio, video and file input, reasoning, tool calling, structured outputs. - `gemini-3.6-flash`: Gemini 3.6 Flash (Google), 1M context (tools, vision, reasoning). $1.013 in / $5.063 out per 1M, cached input $0.10125 per 1M. Google's July 2026 Flash tier. 1M context, accepts text, image, audio and video, and costs less per output token than 3.5 Flash. - `gemini-3.7-flash`: Gemini 3.7 Flash (Google), 1M context (tools, vision, reasoning). $1.013 in / $5.063 out per 1M, cached input $0.050625 per 1M. Google's newest Flash tier. 1M context and 64K max output, the widest input set in the family — text, image, audio, video and file — with reasoning, tool calling and structured outputs. - `kimi-k2.7-code`: Kimi K2.7 Code (Moonshot AI), 262K context (tools, vision, reasoning). $1.009 in / $4.774 out per 1M. Coding specialist. +21.8% Kimi Code Bench vs K2.6, ~30% fewer reasoning tokens. Always-thinking. 262K context. - `hermes-3-70b`: Hermes 3 70B (Nousresearch), 131K context (tools). $0.945 in / $0.945 out per 1M. The 70B of the same tune, at 30% less per million. Keeps the steerability the 405B is chosen for and gives up depth on the hardest reasoning, which is the trade most persona and agent workloads are happy to make. - `kimi-k2.6`: Kimi K2.6 (Moonshot AI), 262K context (tools, vision, reasoning). $0.7856 in / $3.667 out per 1M. Moonshot's newest. Agentic + vision + reasoning. 262K context. - `deepseek-r1`: DeepSeek R1 (DeepSeek), 64K context (tools, reasoning). $0.7425 in / $2.957 out per 1M. Top reasoning model. 96% cheaper than o1. - `deepseek-v4-pro`: DeepSeek V4 Pro (DeepSeek), 1M context (tools, reasoning). $0.6901 in / $1.38 out per 1M, cached input $0.133 per 1M. 1.6T MoE flagship. 1M context. Top reasoning tier. - `gemini-3-flash`: Gemini 3 Flash (Google), 1M context (tools, vision, reasoning). $0.675 in / $4.05 out per 1M, cached input $0.0675 per 1M. Newest Gemini. 1M context. - `nemotron-3-ultra-550b`: Nemotron 3 Ultra 550B (NVIDIA), 1M context (tools, reasoning). $0.675 in / $3.375 out per 1M, cached input $0.135 per 1M. NVIDIA's strongest US open-weight. 550B MoE (55B active), hybrid Mamba-Transformer. 1M context, 300+ tok/s. - `kimi-k2.5`: Kimi K2.5 (Moonshot AI), 262K context (tools, vision, reasoning). $0.6075 in / $3.038 out per 1M. Multimodal agentic. 262K context. - `qwen3.8-27b`: Qwen 3.8 27B (Alibaba), 1M context (tools, vision, reasoning). $0.567 in / $4.05 out per 1M. Alibaba's open-weight (Apache-2.0) dense 27B vision-language model from the Qwen 3.8 line. 1M context, text, image and video input, reasoning, tool calling and structured outputs — the open checkpoint next to the closed Qwen 3.8 Max and Flash. - `qwen-3.6-plus`: Qwen 3.6 Plus (Alibaba), 1M context (tools, vision, reasoning). $0.4911 in / $2.947 out per 1M, prompt caching supported (cached rate not published). Alibaba's newest flagship. #1 on Kyma. - `qwen-3.7-plus`: Qwen 3.7 Plus (Alibaba), 1M context (tools, vision, reasoning). $0.4431 in / $1.773 out per 1M, cached input $0.0864 per 1M. Alibaba's newest Plus flagship. 1M context, vision input, top agentic + reasoning. - `muse-glimmer-30b`: Muse Glimmer 30B (Meta), 131K context (tools, vision, reasoning). $0.405 in / $1.62 out per 1M. Meta's open-weight (Apache 2.0) 30B dense multimodal model, distilled from Muse Spark and tuned for local agentic work. Vision + reasoning, single-GPU friendly. - `gemini-3.5-flash-lite`: Gemini 3.5 Flash Lite (Google), 1M context (tools, vision, reasoning). $0.405 in / $3.375 out per 1M, cached input $0.0405 per 1M. Google's cheapest 1M-context tier. Takes text, image, audio and video; output price includes thinking tokens. - `gemini-2.5-flash`: Gemini 2.5 Flash (Google), 1M context (tools, vision, reasoning). $0.405 in / $3.375 out per 1M, cached input $0.0405 per 1M. Gemini 2.5 Flash. 1M context. RETIRING 2026-10-16 — Google discontinues Gemini 2.5 on AI Studio (Vertex ~10-16→10-20); migrate to gemini-3-flash / gemini-3.6-flash. - `minimax-m2.7`: MiniMax M2.7 (MiniMax), 205K context (tools, reasoning). $0.405 in / $1.62 out per 1M. Next-gen agentic productivity. - `minimax-m3`: MiniMax M3 (MiniMax), 1M context (tools, vision, reasoning). $0.3985 in / $1.594 out per 1M, cached input $0.0756 per 1M. MSA sparse attention. SWE-Bench Pro 59%, Terminal-Bench 66%. Agentic coding, 1M context, multimodal input. - `minimax-m2.5`: MiniMax M2.5 (MiniMax), 197K context (tools, reasoning). $0.3826 in / $1.346 out per 1M, prompt caching supported (cached rate not published). SWE-bench 80.2%. Top agentic coding. - `deepseek-v3`: DeepSeek V3 (DeepSeek), 160K context (tools, reasoning). $0.351 in / $0.513 out per 1M. Previous-gen flagship. Stable, proven. - `qwen-3-coder`: Qwen 3 Coder (Alibaba), 131K context (tools). $0.334 in / $1.519 out per 1M, cached input $0.135 per 1M. Purpose-built for code generation. - `deepseek-v4-flash-vision-exp`: DeepSeek V4 Flash Vision (DeepSeek), 1M context (tools, vision, reasoning). $0.297 in / $0.891 out per 1M. Experimental vision-enabled release of DeepSeek V4 Flash (open weights, MIT, 305B MoE) — the same efficiency tier with image input. DeepSeek labels it experimental; expect the id to move when a stable release lands. - `llama-4-maverick`: Llama 4 Maverick (Meta), 1M context (tools, vision). $0.27 in / $1.08 out per 1M. Meta's larger Llama 4 tier, above Scout. 1M context, vision, served from two independent providers. - `gpt-5.6-luna-pro`: GPT-5.6 Luna Pro (OpenAI), 1M context (tools, vision, reasoning). $0.27 in / $1.62 out per 1M, cached input $0.027 per 1M. Luna at the same list price with the Pro serving profile. 1.05M context, vision, file input; the rate doubles above 272K prompt tokens. - `gpt-5.6-luna`: GPT-5.6 Luna (OpenAI), 1M context (tools, vision, reasoning). $0.27 in / $1.62 out per 1M, cached input $0.027 per 1M. OpenAI's cheapest 5.6 tier. 1.05M context, vision, file input; the rate doubles above 272K prompt tokens. - `step-3.7-flash`: Step 3.7 Flash (StepFun), 256K context (tools, vision, reasoning). $0.27 in / $1.553 out per 1M, cached input $0.054 per 1M. StepFun's fast flash tier. 256K context, multimodal input, tool calling. Cheap throughput. - `qwen3.8-flash`: Qwen 3.8 Flash (Alibaba), 1M context (tools, vision, reasoning). $0.203 in / $0.635 out per 1M, prompt caching supported (cached rate not published). Alibaba's newest Flash tier — ultra cost-efficient multimodal MoE. 1M context, vision, cheaper per token than the Max line. - `glm-4.5-air`: GLM 4.5 Air (Zhipu AI), 131K context (tools, reasoning). $0.1945 in / $1.272 out per 1M, cached input $0.03375 per 1M. Cheap agentic MoE (106B/12B active). - `mimo-v2.5`: MiMo V2.5 (Xiaomi), 1M context (tools, vision, reasoning). $0.189 in / $0.378 out per 1M. Xiaomi's open-weight (MIT) flagship — multimodal, 1M context, reasoning. Strong agentic and coding work at low cost. - `deepseek-v4-flash`: DeepSeek V4 Flash (DeepSeek), 1M context (tools, reasoning). $0.1389 in / $0.2778 out per 1M, cached input $0.0243 per 1M. 284B MoE. 1M context. Fast + cheap V4 tier. - `llama-3.3-70b`: Llama 3.3 70B (Meta), 128K context (tools). $0.135 in / $0.432 out per 1M, prompt caching supported (cached rate not published). Most popular open model. Great all-rounder. - `hy3`: Hunyuan 3 (Tencent), 262K context (tools, reasoning). $0.108 in / $0.446 out per 1M. Tencent's Hunyuan 3 flagship. 256K context, reasoning, tool use. - `qwen-3-32b`: Qwen 3 32B (Alibaba), 33K context (tools, reasoning). $0.108 in / $0.378 out per 1M, prompt caching supported (cached rate not published). Coding-focused 32B. Strong on code, math and multilingual. - `glm-5.3-flash`: GLM 5.3 Flash (Zhipu AI), 1M context (tools, vision, reasoning). $0.101 in / $0.338 out per 1M, prompt caching supported (cached rate not published). Zhipu's newest Flash — MIT-licensed, multimodal, latency/cost-optimized successor to GLM 4.7 Flash. ~1.3M context, vision. - `glm-4.7-flash`: GLM 4.7 Flash (Zhipu AI), 203K context (tools, reasoning). $0.081 in / $0.54 out per 1M, cached input $0.0135 per 1M. Ultra cheap. 200K context. Fast. - `gemma-4-31b`: Gemma 4 31B (Google), 128K context (tools, vision, reasoning). $0.0763 in / $0.218 out per 1M, cached input $0.135 per 1M. Google's newest open model. Multimodal. - `qwen3.7-flash`: Qwen 3.7 Flash (Alibaba), 1M context (tools, vision, reasoning). $0.0498 in / $0.2164 out per 1M, cached input $0.0081 per 1M. Alibaba's vision-language reasoning model. 1M context, takes text, image and video, and is the cheapest model in the catalogue by a wide margin. - `gpt-oss-120b`: GPT-OSS 120B (OpenAI), 128K context (tools, reasoning). $0.0494 in / $0.2406 out per 1M, prompt caching supported (cached rate not published). OpenAI's open source. 120B parameters. ### Retrieval models (per 1M tokens) Embeddings and reranking. These do NOT use chat completions — see the endpoint on each model's page. - `embeddinggemma-300m`: EmbeddingGemma 300M (Google), 2K context. $0.0027 in / $0.00 out per 1M. 768-dimension embeddings from a 300M-parameter model. Built for on-device and high-volume indexing — the cheapest way to embed a large corpus, at a fraction of the usual per-million rate. - `bge-base-en`: BGE Base EN v1.5 (Baai), 1K context. $0.00675 in / $0.00 out per 1M. 768-dimension English embeddings — half the vector of bge-large-en at the same per-million rate, so an index built on it is half the storage and half the comparison cost. Retrieval quality gives up little on short chunks, which is what the 512-token window enforces anyway. - `all-minilm-l12`: all-MiniLM-L12-v2 (Sentence-transformers), 1K context. $0.00675 in / $0.00 out per 1M. 384-dimension embeddings — the smallest vector on the shelf, and the default of the sentence-transformers library, so an enormous amount of existing code expects exactly this shape. Twelve layers where the L6 has six: better recall, still tiny. - `all-minilm-l6`: all-MiniLM-L6-v2 (Sentence-transformers), 1K context. $0.00675 in / $0.00 out per 1M. 384 dimensions from six layers — the most downloaded embedding model there is, and the one most tutorials and starter repos hard-code. Reach for it to match an existing index or to keep a local prototype and a hosted one on the same vectors. - `gte-base`: GTE Base (Alibaba), 1K context. $0.00675 in / $0.00 out per 1M. 768-dimension general text embeddings from Alibaba's Tongyi lab, trained on a broader mixture than BGE and often a point or two ahead of it on out-of-domain retrieval. Same size and same price as bge-base-en, so the choice is which one your corpus likes. - `qwen3-embedding-8b`: Qwen3 Embedding 8B (Alibaba), 33K context. $0.0135 in / $0.00 out per 1M. 4096-dimension embeddings with a 32K input window and strong multilingual retrieval. Use when recall quality matters more than the per-million rate, or when documents are long enough that a 2K window would have to chunk them. - `bge-m3`: BGE-M3 (Baai), 8K context. $0.0135 in / $0.00 out per 1M. 1024-dimension embeddings across 100+ languages, built for retrieval that has to work in more than English. The 8K window takes a whole page without chunking. Reach for it when the corpus is multilingual and the per-million rate matters. - `qwen3-embedding-0.6b`: Qwen3 Embedding 0.6B (Alibaba), 33K context. $0.0135 in / $0.00 out per 1M. 1024-dimension embeddings from the smallest of the Qwen3 retrieval line, at the same rate as its larger siblings. The 32K window is the reason to reach for it: long documents embed whole while the vector stays small enough to index cheaply. - `bge-large-en`: BGE Large EN v1.5 (Baai), 1K context. $0.0135 in / $0.00 out per 1M. 1024-dimension English embeddings, and still one of the most widely benchmarked retrieval models in production. The 512-token window forces chunking, which is why the newer entries above exist — but an index already built on it stays comparable. - `multilingual-e5-large`: Multilingual E5 Large (Microsoft), 1K context. $0.0135 in / $0.00 out per 1M. 1024-dimension embeddings trained on 94 languages, and the most-cited multilingual retrieval baseline outside BGE. Pick it against bge-m3 when the corpus is short-chunk and the comparison is cross-language recall rather than window size. - `qwen3-embedding-4b`: Qwen3 Embedding 4B (Alibaba), 33K context. $0.027 in / $0.00 out per 1M. 2560-dimension embeddings — the middle of the Qwen3 retrieval line. Recall lands between the 0.6B and the 8B, and so does the price. Use it when the 0.6B misses too much and the 8B's 4096-dimension vectors cost too much to store. - `qwen3-reranker-8b`: Qwen3 Reranker 8B (Alibaba), 41K context. $0.135 in / $0.00 out per 1M. Scores how well each document answers a query, so a retriever's top 50 can be cut to the 5 worth sending to a model. Sits between search and synthesis: embeddings decide what to fetch, this decides what survives. A 40K window means whole documents can be judged without chunking. ### Image generation - `minimax-image-01`: MiniMax Image 01 (MiniMax). $0.005 / image. Sub-cent image generation. Cheapest tier on Kyma — $0.005 per image flat regardless of resolution. 5 aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4). Best for high-volume / budget workflows. - `flux-2-pro`: FLUX.2 Pro (Black Forest Labs). $0.0405 / image. BFL's 32B flagship (3× larger than Flux 1.1). Photoreal, multi-reference (up to 10 sources), unified gen+edit, ~60% accurate text-in-image. $0.03/MP base + $0.015 per extra MP. - `nano-banana`: Nano Banana (Google). $0.053 / image. Google Gemini image-gen. Native edit-mode (image-in + prompt → image-out). 3 size tiers (512/1K/2K). RETIRING 2026-10-20 — Google discontinues the Gemini 2.5 endpoint it runs on; migrate to nano-banana-3-flash. - `flux-kontext-pro`: FLUX.1 Kontext Pro (Black Forest Labs). $0.054 / image. Image-to-image edit and refinement. Mask + inpaint. - `recraft-v4`: Recraft V4 (Recraft). $0.054 / image. Top of HF Arena (#1, beats Midjourney V8 / DALL-E 3 / FLUX). Design-aware composition, lighting, textures. - `recraft-v3`: Recraft V3 (Recraft). $0.054 / image. Legacy. Recommend recraft-v4 for new projects (same price, top of HF Arena). - `nano-banana-3-flash`: Nano Banana 3 Flash (Google). $0.061 / image. Newer Gemini 3.1 image-gen. Same edit-mode contract as Nano Banana; sharper output. The migration target for nano-banana, which Google retires 2026-10-20. - `gpt-image-2`: GPT Image 2 (OpenAI). $0.072 / image. OpenAI's flagship image model (Apr 2026). Near-perfect text-in-image (multilingual), reasoning-augmented composition, photorealism, logo-grade lettering. Quality tiers low/medium/high — picker default medium. 1024×1024 / 1024×1536 / 1536×1024 / 2048×2048. - `flux-1.1-ultra`: FLUX 1.1 Pro Ultra (Black Forest Labs). $0.081 / image. Legacy. Recommend flux-2-pro for new projects (cheaper at 1MP, higher quality, multi-reference). - `ideogram-v3`: Ideogram V3 (Ideogram). $0.081 / image. Text-in-image specialist. Best for typography, packaging, logos. - `recraft-v4-vector`: Recraft V4 Vector (Recraft). $0.108 / image. Native SVG output — actual paths + structured layers, edit in Figma/Illustrator. Only model on the market that ships true vector files. - `recraft-v4-pro`: Recraft V4 Pro (Recraft). $0.3375 / image. Recraft V4 at 4MP for print-ready / large-scale assets. Same design taste as V4, higher resolution. - `recraft-v4-vector-pro`: Recraft V4 Vector Pro (Recraft). $0.405 / image. Native SVG at 4MP for print-ready vector assets. Same as V4 Vector with higher detail / scale. ### Video generation - `kling-2.5-pro`: Kling 2.5 Pro (Kuaishou). $0.0945 / second. Cinematic 5-10s video. Photoreal humans, smooth motion. Cheapest Kling tier. T2V or I2V via image_url. - `kling-3-pro`: Kling 3 Pro (Kuaishou). $0.1512 / second. Flagship Kling. Photoreal humans, smooth motion, sharper than 2.5. T2V or I2V via image_url. For native audio, use kling-3-pro-audio. - `kling-3-pro-audio`: Kling 3 Pro (Audio) (Kuaishou). $0.2268 / second. Kling 3 Pro with native audio (ambient + dialogue). Same visuals as kling-3-pro plus synchronized sound. ~50% premium for audio. - `seedance-2-pro`: Seedance 2 Pro (ByteDance). $0.40959 / second. ByteDance flagship video. Multi-shot, native audio bundled, dynamic camera moves. T2V or I2V via image_url. 720p. - `seedance-2-fast`: Seedance 2 Fast (ByteDance). $0.326565 / second. Seedance 2 fast tier — quicker generation, ~20% cheaper than Pro. Native audio bundled. Best for short social clips. - `veo-3-fast`: Veo 3 Fast (Google). $0.135 / second. Google Veo 3 fast tier — 720p, no audio. Cheapest Veo. Balanced quality and speed for social/drafts. - `veo-3`: Veo 3 (Google). $0.54 / second. Google Veo 3 flagship — 1080p with native audio (dialogue + ambient + lip-sync). Top-quality cinematic clips. - `elevenlabs-music`: ElevenLabs Music (ElevenLabs). $0.135 / second. Prompt-driven music generation. Lyrics, instrumental, configurable duration up to 5 min. - `hailuo-02-512p`: Hailuo 02 (512p) (MiniMax). $0.14 / clip. MiniMax Hailuo 02 at 512p — cheapest video tier on Kyma. Flat $0.140 per clip (6s or 10s). T2V or I2V via image_url. Best for social shorts, rapid iteration, budget motion. - `hailuo-02-768p`: Hailuo 02 (768p) (MiniMax). $0.42 / clip. Hailuo 02 at 768p — mid tier balanced quality vs cost. Flat $0.420 per clip. T2V or I2V via image_url. - `hailuo-02-1080p`: Hailuo 02 (1080p) (MiniMax). $0.78 / clip. Hailuo 02 at 1080p — premium tier, full HD output. Flat $0.780 per clip. T2V or I2V via image_url. ### Audio: speech, transcription, music, sound - `whisper-v3-turbo`: Whisper Large v3 Turbo (OpenAI). $0.0009 / minute. Speech-to-text. 228x realtime inference. Transcripts with timestamps + language detect. - `gpt-4o-mini-transcribe-2025-12-15`: GPT-4o mini Transcribe (OpenAI). $0.00405 / minute. Speech-to-text. OpenAI's premium quality STT — best real-world accuracy on conversational audio, noisy backgrounds, and code-switching (Vi/En etc). RETIRING 2027-02-26 — OpenAI shuts down the gpt-4o-transcribe family; migrate to gpt-transcribe. - `gpt-transcribe`: GPT Transcribe (OpenAI). $0.006075 / minute. Speech-to-text. OpenAI's current premium STT and the replacement for the gpt-4o-transcribe family — high real-world accuracy on conversational audio, noisy backgrounds, and code-switching (Vi/En etc). - `gemini-3.5-transcribe`: Gemini 3.5 Transcribe (Google). $0.00675 / minute. Speech-to-text. Google's dedicated file transcription model — 85+ languages, up to 1 hour per request. OpenAI-compatible /v1/audio/transcriptions. - `gemini-3-flash-audio`: Gemini 3 Flash (Audio) (Google). $0.002592 / minute. Audio understanding. Hears tone, music, SFX, language, speaker emotion — beyond pure transcription. Inline payload up to 30 min. - `gpt-realtime-translate`: GPT Realtime Translate (OpenAI). $0.0459 / minute. Native audio-to-audio translation with voice cloning. Preserves original speaker tone. 13 target languages (es/pt/fr/ja/ru/zh/de/ko/hi/id/vi/it/en). - `gemini-2.5-flash-native-audio-preview-12-2025`: Gemini 2.5 Flash Native Audio (Google). $0.03888 / minute. Conversational realtime with 30 pickable voices + 24 output languages. WebSocket-based, ephemeral token auth. Native audio understanding + generation in one round-trip. RETIRING 2026-10-20 — Google discontinues the Gemini 2.5 endpoint it runs on; gemini-3.1-flash-live-preview is the successor at the same per-minute price. - `gemini-3.1-flash-live-preview`: Gemini 3.1 Flash Live (Google). $0.03888 / minute. Audio-to-audio realtime dialogue on the Gemini 3 line, at the same per-minute price as Gemini 2.5 Flash Native Audio. The successor to that model, which retires 2026-10-20. WebSocket-based; native audio in and out in one round-trip. - `gemini-3.5-live-translate-preview`: Gemini 3.5 Live Translate (Google). $0.06345 / minute. Low-latency audio-to-audio speech translation. Near real-time speech-to-speech across 70+ languages, preserving the speaker's intonation, pacing, and pitch. WebSocket-based realtime session. - `eleven-multilingual-v2`: ElevenLabs Multilingual v2 (ElevenLabs). $0.405 / 1K characters. Hero-quality multilingual TTS. 29 languages, expressive voices, brand-safe consistent delivery. - `eleven-v3`: ElevenLabs v3 (ElevenLabs). $0.405 / 1K characters. Most expressive TTS. Emotional range, audio tags, and lifelike delivery across 70+ languages. - `eleven-flash-v2-5`: ElevenLabs Flash v2.5 (ElevenLabs). $0.2025 / 1K characters. Ultra-low-latency TTS, ~75ms time-to-first-byte. Half the per-char cost of Multilingual v2. 32 languages. - `eleven-turbo-v2-5`: ElevenLabs Turbo v2.5 (ElevenLabs). $0.2025 / 1K characters. Balanced TTS — quicker than Multilingual, better quality than Flash. Half cost vs Multilingual. 32 languages. - `elevenlabs-sfx`: ElevenLabs Sound Effects (ElevenLabs). $0.027 / generation. Generates non-speech audio (whoosh, explosion, rain) from a text prompt. Flat $0.027 per generation, 0.5-22 sec. - `minimax-speech-hd`: MiniMax Speech HD (MiniMax). $0.07 / 1K characters. MiniMax HD voice. Multilingual, expressive, ~2.9× cheaper than ElevenLabs Multilingual v2 at the same production quality tier. - `minimax-speech-turbo`: MiniMax Speech Turbo (MiniMax). $0.04 / 1K characters. MiniMax low-latency voice. Multilingual, ~2.2× cheaper than ElevenLabs Flash v2.5. Best for bulk TTS, real-time voice agents, conversational AI. - `minimax-music`: MiniMax Music (MiniMax). $0.20 / track. Lyrics-driven music generation. Music-2.0 family. Up to 5 minutes per call, ~90× cheaper than ElevenLabs Music for non-hero use cases. - `minimax-music-pro`: MiniMax Music Pro (MiniMax). $0.20 / track. Music-2.6 (latest pro family). Higher fidelity than Music-2.0, richer arrangements. Still ~19× cheaper than ElevenLabs Music for production-tier output. - `minimax-voice-clone`: MiniMax Voice Clone (MiniMax). $2.00 / generation. Clone a voice from a 10s-5min reference recording. Returns a voice_id usable in /v1/audio/speech with any MiniMax HD/Turbo SKU. Flat one-time charge per cloned voice. - `minimax-voice-design`: MiniMax Voice Design (MiniMax). $2.00 / generation. Generate a synthesized voice profile from a natural-language description (no reference audio needed). Returns a voice_id usable in /v1/audio/speech with any MiniMax HD/Turbo SKU. Flat one-time charge per designed voice. ## Model Aliases Send these as the model name and Kyma resolves to the best current model: - `best` → `qwen-3.6-plus` - `fast` → `gemini-3.5-flash-lite` - `code` → `qwen-3-coder` - `cheap` → `deepseek-v4-flash` - `long-context` → `gemini-3-flash` - `vision` → `gemma-4-31b` - `reasoning` → `deepseek-r1` - `agent` → `kimi-k2.6` - `best-agent` → `kimi-k2.6` - `balanced` → `llama-3.3-70b` - `glm-flagship` → `glm-5.2` - `search` → `sonar` - `transcribe` → `whisper-v3-turbo` - `transcribe-quality` → `gpt-4o-mini-transcribe-2025-12-15` - `audio-understand` → `gemini-3-flash-audio` ## Key Facts - OpenAI SDK compatible: set base_url to https://kymaapi.com/v1 and use a ky- API key - Anthropic Messages API compatible: POST /v1/messages - Image generation: POST /v1/images/generations (async, poll GET /v1/jobs/{id}) - Video generation: POST /v1/videos/generations (async, poll GET /v1/jobs/{id}) - Audio: transcription (POST /v1/audio/transcriptions), text-to-speech (POST /v1/audio/speech), understanding (POST /v1/audio/understand) - Multi-route redundancy: automatic failover keeps user-facing errors near zero - Prompt caching: cached input billed at each model's own cached rate, automatic where supported - Tool calling and JSON structured outputs: supported on language models - Rate limits: tier-based, scaling with credits purchased - Free tier: the $0.50 signup credit spends on 43 of 126 models: everything billed per token (language, embedding, reranking) at or under $3/M output and $0.50/M input, plus everything billed per unit at or under $0.01 per billing unit. Everything else that generates media needs a top-up; free-tier media today: `gemini-3-flash-audio`, `gemini-3.5-transcribe`, `gpt-4o-mini-transcribe-2025-12-15`, `gpt-transcribe`, `minimax-image-01`, `whisper-v3-turbo` - Signup: https://kymaapi.com or POST /v1/auth/register ## Optional - [Changelog](https://docs.kymaapi.com/changelog): Recent updates - [Use Cases](https://docs.kymaapi.com/guides/use-cases/chatbot): Chatbot, coding agent, RAG, data extraction, automation - [MCP Server](https://docs.kymaapi.com/guides/mcp-server): Streamable HTTP at https://mcp.kymaapi.com/mcp (OAuth) plus local/dev stdio