Alibaba

Alibaba

Qwen3 Embedding 8B

Qwen3 Embedding 8B generates 4096-dimensional text embeddings optimized for multilingual retrieval and long-context RAG pipelines. Reach for it when document length or cross-language recall outweighs the need for the lowest possible per-token rate.

Modalities

Text → Text

Usage

How much this model is actually called here.

Rank

#65

of 98 active models

Tokens served

118.9K

all-time

Platform share

0.0%

of all tokens

Tokens · last 15 daysSep 1Sep 15

Pricing

Pay per token. Cached input bills at this model’s own cached rate, listed below.

$0.0135 /1M input$0.00 /1M output
Hobby10 req/day · 2K in / 500 out
~$0.01/mo
Production1,000 req/day · 2K in / 500 out
~$0.81/mo
Scale20,000 req/day · 2K in / 500 out
~$16.20/mo
+ Estimate your workload
1,000
2,000
500

Estimated monthly cost

$0.8100

$0.0270 / day on Qwen3 Embedding 8B

Same workload on:

Qwen 3.8 27B$94.77+11600%
Qwen 3.7 Flash$6.23+670%

Estimates use list pricing. Actual bills depend on real token counts, and every response includes its exact cost.

When to use Qwen3 Embedding 8B

Updated 2026-07-31

Where this model earns its cost — and where it doesn't.

This model produces dense vector representations with a 32,768-token context window. It is designed for text-only embedding tasks and does not support reasoning, vision, or structured output generation.

On Kyma, it runs as an OpenAI-compatible endpoint behind a single API key. Requests benefit from automatic failover if a serving path degrades, and responses return the exact cost in usage.cost alongside an X-Kyma-Model header. Prompt caching is supported, billing repeated prefixes at this model's cached input rate.

Because it is strictly an embedding model, it returns zero output tokens and cannot generate conversational text. It operates at a medium speed tier, making it better suited for batch indexing or retrieval-heavy workflows than for real-time, low-latency interactive search.

Multilingual Document Search

Finds relevant passages across different languages without translation overhead.

Long-Form Context Indexing

Embeds full documents up to 32K tokens to avoid aggressive chunking.

High-Recall Retrieval Pipelines

Prioritizes semantic match quality over minimal compute cost for RAG systems.

Not ideal for: Do not use this model for real-time chat, text generation, or tasks requiring sub-100ms latency, as it only outputs vectors and runs at a medium speed tier.

How it compares

Against the peers people actually weigh it against.

SpecQwen3 Embedding 8BQwen 3.8 27BQwen 3.7 Flash
Input /1M$0.0135$0.567$0.0498
Output /1M$0.00$4.05$0.2164
Context33K1M1M
ToolsNoYesYes
ReasoningNoYesYes
Throughput252.8 tok/s45.6 tok/s117.8 tok/s

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit on the free tier, which covers Qwen3 Embedding 8B. No card required.

    Get API key →
  2. 2

    Make your first request

    Submit a generation job and poll until it succeeds.

    curl https://kymaapi.com/v1/embeddings \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen3-embedding-8b",
        "input": ["first document", "second document"]
      }'

FAQ

Common questions about this model.

What is the context window of Qwen3 Embedding 8B?

Qwen3 Embedding 8B has a 33K-token context window — roughly 48 pages of text in a single request.

How much does the Qwen3 Embedding 8B API cost?

$0.0135 per 1M input tokens and $0.00 per 1M output tokens. No subscription; you pay only for what you use.

Does Qwen3 Embedding 8B support function calling?

No — Qwen3 Embedding 8B does not support tool calling. For agent workloads, choose a tools-enabled model from the catalog.

How do I use Qwen3 Embedding 8B?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to qwen3-embedding-8b. Signing up is free and includes $0.50 of credit on the free tier, which covers this model — no card required.

Does this model generate text responses?

No, it is strictly an embedding model that outputs 4096-dimensional vectors and returns zero output tokens.

How does Kyma handle prompt caching for this endpoint?

Kyma supports caching for embedding requests, which have no reusable prompt prefix and therefore no cached rate.

Can I use the same API key for this model as I do for others?

Yes, Kyma is OpenAI-compatible and uses a single API key across all models at the standard base URL.

Start with $0.50 free credit on the free tier — no card required.Create account →

More models by Alibaba

See all 14
ModelContextInputOutput
AlibabaQwen 3.8 Flash1M$0.203$0.635
AlibabaQwen 3.8 27B1M$0.567$4.05
AlibabaQwen 3.8 Max1M$2.2275$6.684
AlibabaQwen 3.7 Flash1M$0.0498$0.2164
AlibabaQwen 3.7 Plus1M$0.4431$1.773
AlibabaQwen 3.7 Max1M$2.304$6.909
AlibabaQwen 3.6 Plus1M$0.4911$2.947
AlibabaQwen 3 Coder131K$0.334$1.519