# Gemini 3 Flash

> Newest Gemini. 1M context.

Human version: https://kymaapi.com/models/gemini-3-flash
Live JSON: `GET https://kymaapi.com/v1/models` (no auth required)

## Facts

- **Model ID**: `gemini-3-flash` — pass this as `model` in the request body
- **Creator**: Google
- **Released**: 2025-12-17
- **Context window**: 1M tokens
- **Max output**: 8K tokens per response — a hard ceiling, not a default
- **Price**: $0.3375 in / $1.35 out per 1M, cached input 10%
- **Capabilities**: tools, vision, reasoning, caching

## Call it

```bash
curl https://kymaapi.com/v1/chat/completions \
  -H "Authorization: Bearer $KYMA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gemini-3-flash", "messages": [{"role": "user", "content": "Hello"}]}'
```

## Positioning

Google's newest Gemini on Kyma, pairing a 1M-token context window with text, image, audio, and video inputs — the model to reach for when a request needs to see, hear, or read a lot at once.

## About

Gemini 3 Flash is the newest Gemini from Google, built for long-context work and reasoning. It sits in Kyma's frontier-open quality tier and accepts text, images, audio, and video in a single request, returning text — with extended reasoning available when the problem calls for it.

On Kyma it's a top-10 model by production traffic — ranked 8th of 62 by tokens served, with a 99.8% success rate across recent requests. Python apps and the OpenClaw coding agent lead its traffic. Every call gets automatic failover if a serving path degrades, and prompt caching bills repeated prompt prefixes at 10% of the input rate, which matters at this context size — resending a large cached prefix costs a fraction of the first pass.

The 1,048,576-token context window comes with function calling and structured outputs, so the long context is usable inside agent pipelines, not just for one-off summarization.

## Use cases

- **Whole-corpus analysis** — The 1M-token window fits entire codebases, document sets, or transcript archives in one request instead of a retrieval pipeline.
- **Video and audio understanding** — Send recordings, screen captures, or audio directly — no separate transcription step — and ask questions about what's in them.
- **Vision tasks** — Screenshots, diagrams, charts, and scanned documents go in as images alongside your text prompt.
- **Reasoning over long inputs** — Extended reasoning plus the huge context handles analysis that requires holding a lot of material in view at once.
- **Long-context agents** — Function calling and structured outputs keep multi-step agents reliable even as the working context grows toward the 1M-token window.

## Not ideal for

Single responses that need to run very long — output is capped at 8K tokens per request, so generating a book-length draft means chunking; and it sits in the premium cost tier, so high-volume simple tasks are cheaper on a lighter model.

## Pick something else when

- You want the safer long-context default → `gemini-2.5-flash`
- You want the best overall default → `qwen-3.6-plus`
- You need the strongest coding agent behavior → `kimi-k2.6`

## See also

- All models: https://kymaapi.com/models.md
- Pricing: https://kymaapi.com/pricing.md
- Other models by Google: https://kymaapi.com/models?q=Google
