EmbeddingGemma 300M is a lightweight, 768-dimension text embedding model optimized for high-volume indexing and budget-constrained retrieval. Reach for it when you need to process large text corporas quickly without paying premium embedding rates.
Modalities
Text → Text
Usage
How much this model is actually called here.
Rank
#77
of 98 active models
Tokens served
50.6K
all-time
Platform share
0.0%
of all tokens
Pricing
Pay per token. Cached input bills at this model’s own cached rate, listed below.
+ Estimate your workload− Estimate your workload
Estimated monthly cost
$0.1620
$0.0054 / day on EmbeddingGemma 300M
Same workload on:
Estimates use list pricing. Actual bills depend on real token counts, and every response includes its exact cost.
When to use EmbeddingGemma 300M
Where this model earns its cost — and where it doesn't.
Google’s EmbeddingGemma 300M generates 768-dimension text embeddings from a compact 300M-parameter architecture. It is designed specifically for bulk corpus indexing, semantic deduplication, and retrieval-augmented generation pipelines where cost efficiency and throughput matter more than maximum representational depth.
On Kyma, the model runs through an OpenAI-compatible endpoint with automatic request failover and exact cost reporting in the usage object. It supports prompt caching, which bills repeated prefixes at this model's cached input rate.
The model accepts text input up to a 2048-token context window and outputs fixed-length vectors. It does not support text generation, reasoning, vision, or structured outputs, and it is strictly an embedding-only endpoint.
Bulk Document Indexing
Index large text corporas efficiently while keeping per-token embedding costs minimal.
Budget RAG Retrieval
Power semantic search in retrieval pipelines where high throughput outweighs the need for larger vector dimensions.
Semantic Text Deduplication
Identify and remove duplicate or near-duplicate entries across large datasets using fast vector similarity.
High-Volume Data Processing
Process massive text batches quickly with a lightweight model optimized for speed and low latency.
Not ideal for: Do not use this model for documents exceeding 2048 tokens, tasks requiring generative text or reasoning, or applications that demand high-dimensional embeddings for fine-grained semantic discrimination.
How it compares
Against the peers people actually weigh it against.
| Spec | EmbeddingGemma 300M | Gemini 3.8 Flash | Gemini 3.5 Flash Lite |
|---|---|---|---|
| Input /1M | $0.0027 | $1.013 | $0.405 |
| Output /1M | $0.00 | $5.063 | $3.375 |
| Context | 2K | 1M | 1M |
| Tools | No | Yes | Yes |
| Reasoning | No | Yes | Yes |
| Throughput | 1210.8 tok/s | 35 tok/s | 53.5 tok/s |
Quick start
Up and running in under two minutes.
- 1
Create an API key
Sign up and grab a key from the dashboard — $0.50 free credit on the free tier, which covers EmbeddingGemma 300M. No card required.
Get API key → - 2
Make your first request
Submit a generation job and poll until it succeeds.
curl https://kymaapi.com/v1/embeddings \ -H "Authorization: Bearer YOUR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "embeddinggemma-300m", "input": ["first document", "second document"] }'
FAQ
Common questions about this model.
What is the context window of EmbeddingGemma 300M?
How much does the EmbeddingGemma 300M API cost?
Does EmbeddingGemma 300M support function calling?
How do I use EmbeddingGemma 300M?
What is the maximum input length for this model?
Does Kyma support prompt caching for this endpoint?
Can I use this model to generate text or answer questions?
More models by Google
See all 14 →| Model | Context | Input | Output |
|---|---|---|---|
| 1M | $1.013 | $5.063 | |
| — | $0.00675 / min | ||
| 1M | $1.013 | $5.063 | |
| 1M | $1.013 | $5.063 | |
| 1M | $0.405 | $3.375 | |
| 1M | $2.025 | $12.15 | |
| 128K | $0.0763 | $0.218 | |
| — | $0.061 / image | ||