Models

126 models, one API key. Open-source text, code, image, and video — pay per token, image, or second.

CompareModelPrecisionCapabilities
1M$1.013
$2.026
$5.063
$10.126
1M$13.50$67.50
$0.0061 / min
$0.0067 / min
1M$0.203$0.63544tok/s2.64s
46.5%
1M$0.101
$0.203
$0.338
$0.675
30tok/s2.50s
91.2%
1M$0.297$0.891149tok/s0.52s
100%
1M$1.89$5.9432tok/s2.32s
96.4%
1M$0.567$4.0550tok/s1.55s
100%
1M$1.013
$2.026
$5.063
$10.126
$0.0506391tok/s2.98s
96.6%
500K$2.70$8.1057tok/s4.76s
97.0%
131K$0.405$1.6257tok/s5.01s
82.1%
1M$1.688$5.738147tok/s3.46s
87.5%
1M$2.2275$6.684$0.278128tok/s2.96s
97.3%
1M$0.189$0.37824tok/s3.27s
91.1%
1M$0.0498$0.2164$0.0081108tok/s2.51s
87.9%
1M$13.50$67.50$1.3533tok/s1.44s
95.4%
1M$6.75$33.75$0.67520tok/s1.90s
89.3%
1M$1.013$5.063$0.1012529tok/s2.50s
98.4%
1M$0.405$3.375$0.040557tok/s1.07s
96.2%
1M$1.542$5.243178tok/s2.02s
94.1%
1M$4.05$20.25$0.40528tok/s2.61s
87.5%
1M$2.70
$5.40
$13.50
$27.00
$0.337552tok/s3.82s
95.5%
1M$2.70
$5.40
$13.50
$27.00
$0.337528tok/s1.57s
97.5%
1M$2.70$16.20$0.2776tok/s2.95s
86.6%
1M$2.70$16.20$0.2737tok/s1.22s
95.8%
1M$0.27$1.62$0.02777tok/s3.00s
92.3%
1M$0.27$1.62$0.02739tok/s1.17s
94.3%
500K$2.73$8.18948tok/s2.48s
95.8%
262K$0.108$0.44671tok/s3.76s
83.9%
1M$2.70$13.50$0.2718tok/s4.25s
96.8%
1M$1.101$3.76$0.29835fp4 · fp878tok/s2.81s
97.5%
262K$1.009$4.77447tok/s1.79s
99.1%
1M$13.50$67.5016tok/s3.59s
96.9%
$0.0635 / min
1M$0.675$3.375$0.135fp4 · fp852tok/s1.50s
98.0%
1M$0.4431$1.773$0.086454tok/s16.79s
60.8%
1M$0.3985$1.594$0.0756fp4 · fp834tok/s2.11s
97.3%
256K$0.27$1.553$0.054fp850tok/s3.00s
97.8%
1M$2.304$6.909$0.1687556tok/s15.80s
79.4%
256K$1.389$2.777$0.27106tok/s2.66s
96.2%
1M$2.025$12.15$0.2025123tok/s2.26s
97.1%
$0.0459 / min
1M$1.92$3.838$0.2786tok/s2.21s
96.1%
1M$0.6901$1.38$0.13332tok/s3.54s
98.1%
1M$0.1389$0.2778$0.024330tok/s1.47s
99.9%
$0.072 / image
98.7%
262K$0.7856$3.66730tok/s1.28s
97.2%
1M$6.75$33.75$0.67524tok/s1.77s
94.7%
$0.4096 / sec
$0.3266 / sec
$0.2 / song
203K$1.89$5.94$0.27675fp4 · fp829tok/s11.59s
99.3%
131K$0.655$3.9351tok/s8.58s
82.1%
128K$0.0763$0.218$0.135fp4 · fp810tok/s5.03s
99.9%
2M$1.767$3.534308tok/s7.50s
97.1%
2M$1.763$3.52839tok/s1.09s
96.2%
$0.0389 / min
205K$0.405$1.6254tok/s5.28s
91.6%
$0.061 / image
1M$2.70$16.20$0.2796tok/s4.84s
94.6%
$0.108 / image
1M$4.05$20.25$0.40531tok/s1.94s
96.5%
$0.054 / image
$0.3375 / image
$0.405 / image
197K$0.3826$1.34669tok/s3.90s
99.8%
$0.1512 / sec
$0.2268 / sec
262K$0.6075$3.03835tok/s5.54s
92.6%
203K$0.081$0.54$0.0135fp861tok/s4.55s
95.7%
1M$0.675$4.05$0.067540tok/s1.42s
98.7%
$0.0026 / min
$0.0389 / min
160K$0.351$0.513fp412tok/s1.20s
83.1%
$0.0405 / image
$0.2 / song
200K$1.35$6.75$0.13539tok/s1.44s
95.6%
$0.0945 / sec
2K$0.0027$0.001182tok/s0.34s
100%
$0.053 / image
$0.135 / sec
128K$0.0494$0.240634tok/s3.58s
92.4%
$0.135 / sec
131K$0.1945$1.272$0.03375fp837tok/s3.10s
96.9%
131K$0.334$1.519$0.135fp4 · fp848tok/s1.18s
99.1%
$2.00 / call
$0.78 / video
$0.42 / video
$0.14 / video
$0.405 / 1K char31ch/s4.06s
33K$0.0135$0.00197tok/s2.04s
99.7%
33K$0.0135$0.001162tok/s0.34s
100%
33K$0.027$0.001069tok/s0.38s
100%
41K$0.135$0.00
$0.054 / image
$0.54 / sec
33K$0.108$0.37846tok/s8.95s
91.4%
1M$0.27$1.08fp833tok/s1.13s
94.0%
$2.00 / call
$0.07 / 1K char39ch/s3.25s
$0.04 / 1K char41ch/s3.06s
100%
$0.081 / image
1M$0.405$3.375$0.040542tok/s1.08s
96.4%
200K$4.05$20.2525tok/s1.75s
95.3%
127K$1.35$1.3523tok/s1.91s
97.2%
$0.005 / image
14.9%
64K$0.7425$2.95766tok/s1.88s
99.1%
128K$0.135$0.43215tok/s0.95s
86.6%
$0.081 / image
$0.054 / image
$0.2025 / 1K char112ch/s1.12s
$0.0009 / min8.4×rt1.20s
100%
131K$1.35$1.35fp817tok/s0.94s
100%
131K$0.945$0.945fp825tok/s0.85s
96.9%
$0.2025 / 1K char110ch/s1.14s
$0.027 / call
$0.004 / min4.9×rt2.03s
100%
1K$0.0135$0.00430tok/s1.19s
100%
8K$0.0135$0.001289tok/s0.41s
97.7%
1K$0.0135$0.001331tok/s0.30s
100%
1K$0.00675$0.001391tok/s0.29s
100%
$0.405 / 1K char56ch/s2.26s
1K$0.00675$0.00442tok/s0.91s
100%
1K$0.00675$0.0011tok/s11.58s
100%
1K$0.00675$0.00285tok/s0.90s
100%

Token models priced per 1M tokens. The cached column is what a repeated prompt prefix costs on that model — a rate of its own, read from the same source as the input rate beside it, not a fixed fraction of it. It runs from a tenth of the input rate to the full input rate depending on the model, and a dash means nobody publishes one, which is a different statement from “free”. Image, video, and audio models priced per unit (image, second, video, minute, song, call, or 1K characters). Throughput and latency are measured by Kyma every six hours with an identical input sent to every model through this API — so the numbers compare. For chat and embedding models throughput is tokens per second; for transcription models it is the realtime factor (×rt: seconds of audio transcribed per second of wall time); for text-to-speech it is characters synthesised per second (ch/s) — each off one fixed input. Image, video, music and realtime models are not probed; their reliability comes from real jobs, so those two columns stay empty by design. Uptime is every observation over 30 days, probes and real requests together. Precision names the numeric formats a model may be served at that are BELOW the one its creator released it in, and it is filled on the 14 models where that is true. A low format is not a downgrade when it is how the model shipped, so this is a comparison against the creator’s own model card rather than a list of small numbers. An empty cell means there is nothing to flag: either no route reports going under the release, or nobody could source what the release was. Those two are different answers and the model’s own page prints which one it is.

Start building with any model

$0.50 free credits on signup. OpenAI-compatible API.

Get Free API Key