# Gemini 3.5 Flash

> Newest Gemini Flash. 1M context, multimodal input.

Human version: https://kymaapi.com/models/gemini-3.5-flash
Live JSON: `GET https://kymaapi.com/v1/models` (no auth required)

## Facts

- **Model ID**: `gemini-3.5-flash` — pass this as `model` in the request body
- **Creator**: Google
- **Released**: 2026-05-19
- **Context window**: 1M tokens
- **Max output**: 8K tokens per response — a hard ceiling, not a default
- **Price**: $2.025 in / $12.15 out per 1M, cached input 10%
- **Capabilities**: tools, vision, reasoning, caching

## Call it

```bash
curl https://kymaapi.com/v1/chat/completions \
  -H "Authorization: Bearer $KYMA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gemini-3.5-flash", "messages": [{"role": "user", "content": "Hello"}]}'
```

## Positioning

Google's newest Gemini Flash on Kyma: a 1M-token context window with text, image, audio, and video input in one fast model. Reach for it when a request needs to see or hear something — or when the whole document has to fit in one request.

## About

Gemini 3.5 Flash is the newest Flash-class model from Google, built for long context, multimodal input, and fast reasoning. It sits in Kyma's frontier-open quality tier and takes text, images, audio, and video as input while returning text.

On Kyma it runs through the same OpenAI-compatible endpoint as every other model, with automatic failover if a serving path degrades. Prompt caching is supported, so repeated prompt prefixes — long system prompts, large documents you query more than once — bill at 10% of the input rate. Every response reports its exact cost in usage.cost.

The 1M-token context window is the headline capability: whole codebases, long transcripts, or large document sets fit in a single request. Function calling, structured outputs, and reasoning support round it out for agent work, not just one-shot prompts.

## Use cases

- **Video and audio analysis** — Feed it video or audio directly — summarize recordings, extract what was said and shown, no separate transcription step.
- **Whole-corpus questions** — The 1M context takes an entire codebase, contract set, or research archive in one request instead of a chunked retrieval pipeline.
- **Vision pipelines** — Image understanding with structured outputs turns screenshots, documents, and photos into clean JSON your code consumes.
- **Fast reasoning at scale** — A fast-tier model that also supports reasoning — multi-step analysis without leaving the fast tier.
- **Multimodal agents** — Tool calling plus four input modalities lets one agent handle text, screenshots, and recordings without switching models.

## Not ideal for

Long single-shot generations — output caps at 8K tokens — or cost-sensitive bulk text work, where its premium pricing buys multimodal range you wouldn't be using.

## Pick something else when

- You want the cheapest long-context option → `gemini-2.5-flash`
- You need the strongest tool-heavy agent behavior → `kimi-k2.6`
- You want top reasoning over speed → `deepseek-v4-pro`

## See also

- All models: https://kymaapi.com/models.md
- Pricing: https://kymaapi.com/pricing.md
- Other models by Google: https://kymaapi.com/models?q=Google
