Sentence Transformers

Sentence Transformers

all-MiniLM-L12-v2

384-dimension embeddings — the smallest vector on the shelf, and the default of the sentence-transformers library, so an enormous amount of existing code expects exactly this shape. Twelve layers where the L6 has six: better recall, still tiny.

Modalities

Text → Text

Usage

How much this model is actually called here.

Rank

#85

of 98 active models

Tokens served

15.9K

all-time

Platform share

0.0%

of all tokens

Tokens · last 15 daysSep 3Sep 17

Pricing

Pay per token. Cached input bills at this model’s own cached rate, listed below.

$0.00675 /1M input$0.00 /1M output
Hobby10 req/day · 2K in / 500 out
~$0.00/mo
Production1,000 req/day · 2K in / 500 out
~$0.40/mo
Scale20,000 req/day · 2K in / 500 out
~$8.10/mo
+ Estimate your workload
1,000
2,000
500

Estimated monthly cost

$0.4050

$0.0135 / day on all-MiniLM-L12-v2

Same workload on:

all-MiniLM-L6-v2$0.4050
EmbeddingGemma 300M$0.1620-60%

Estimates use list pricing. Actual bills depend on real token counts, and every response includes its exact cost.

How it compares

Against the peers people actually weigh it against.

Specall-MiniLM-L12-v2all-MiniLM-L6-v2EmbeddingGemma 300M
Input /1M$0.00675$0.00675$0.0027
Output /1M$0.00$0.00$0.00
Context1K1K2K
ToolsNoNoNo
ReasoningNoNoNo
Throughput10 tok/s259 tok/s1211 tok/s

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit on the free tier, which covers all-MiniLM-L12-v2. No card required.

    Get API key →
  2. 2

    Make your first request

    Submit a generation job and poll until it succeeds.

    curl https://kymaapi.com/v1/embeddings \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "all-minilm-l12",
        "input": ["first document", "second document"]
      }'

FAQ

Common questions about this model.

What is the context window of all-MiniLM-L12-v2?

all-MiniLM-L12-v2 has a 1K-token context window — roughly 1 pages of text in a single request.

How much does the all-MiniLM-L12-v2 API cost?

$0.00675 per 1M input tokens and $0.00 per 1M output tokens. No subscription; you pay only for what you use.

Does all-MiniLM-L12-v2 support function calling?

No — all-MiniLM-L12-v2 does not support tool calling. For agent workloads, choose a tools-enabled model from the catalog.

How do I use all-MiniLM-L12-v2?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to all-minilm-l12. Signing up is free and includes $0.50 of credit on the free tier, which covers this model — no card required.

Start with $0.50 free credit on the free tier — no card required.Create account →

More models by Sentence Transformers

ModelContextInputOutput
Sentence Transformersall-MiniLM-L6-v21K$0.00675$0.00