Alibaba

Alibaba

Qwen 3.8 Max

Alibaba's newest flagship, GA successor to the 3.7 Max line — and the first Max with vision. 1M context, multimodal reasoning over text and images, cheaper per token than the 3.7 it replaces.

Modalities

Text+Image → Text

Input

$2.2275 /1M

Output

$6.684 /1M

Cached input

$0.22275 /1M90% off

Context

1M

Speed

medium

Pricing

Pay per token. Cached input is billed at 10% of the input rate.

$2.2275 /1M input$6.684 /1M output
Hobby10 req/day · 2K in / 500 out
~$2.34/mo
Production1,000 req/day · 2K in / 500 out
~$234/mo
Scale20,000 req/day · 2K in / 500 out
~$4,678/mo
+ Estimate your workload
1,000
2,000
500
30%

Estimated monthly cost

$198

$6.59 / day on Qwen 3.8 Max

Same workload on:

Qwen 3.7 Flash$5.73-97%
Qwen 3.7 Max$269+36%

Estimates use list pricing with cached input billed at the 90%-discount rate. Actual bills depend on real token counts — every response includes its exact cost.

How it compares

Against the peers people actually weigh it against.

SpecQwen 3.8 MaxQwen 3.7 FlashQwen 3.7 Max
Input /1M$2.2275$0.0458$2.56
Output /1M$6.684$0.1987$7.676
Context1M1M1M
ToolsYesYesYes
ReasoningYesYesYes
Speedmediumfastmedium

Quick start

Up and running in under two minutes.

  1. 1

    Create an API key

    Sign up and grab a key from the dashboard — $0.50 free credit, no card required.

    Get API key →
  2. 2

    Make your first request

    Drop in your key and send a chat completion — fully OpenAI-compatible.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen-3.8-max",
        "messages": [
          {"role": "user", "content": "Explain prompt caching in one paragraph."}
        ]
      }'
  3. 3

    Stream responses

    Add "stream": true to receive tokens as they arrive.

    curl https://kymaapi.com/v1/chat/completions \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "qwen-3.8-max",
        "stream": true,
        "messages": [{"role": "user", "content": "Hello!"}]
      }'

FAQ

Common questions about this model.

What is the context window of Qwen 3.8 Max?

Qwen 3.8 Max has a 1M-token context window — roughly 1471 pages of text in a single request.

How much does the Qwen 3.8 Max API cost?

$2.2275 per 1M input tokens and $6.684 per 1M output tokens, with cached input at $0.22275/1M — a 90% discount on repeated prompt prefixes. No subscription; you pay only for what you use.

Does Qwen 3.8 Max support function calling?

Yes — Qwen 3.8 Max supports tool/function calling and structured outputs (JSON mode), so it works with agent frameworks out of the box.

How do I use Qwen 3.8 Max?

Kyma is OpenAI-compatible: point your SDK's base URL at https://kymaapi.com/v1, use your Kyma API key, and set the model to qwen-3.8-max. Signing up is free and includes $0.50 of credit — no card required.

Start with $0.50 free credit — no card required.Create account →

More models by Alibaba

See all 12
ModelContextInputOutput
AlibabaQwen 3.7 Flash1M$0.0458$0.1987
AlibabaQwen 3.7 Plus1M$0.4482$1.793
AlibabaQwen 3.7 Max1M$2.56$7.676
AlibabaQwen 3.6 Plus131K$0.4388$2.633
AlibabaQwen 3 Coder131K$0.297$1.35
AlibabaQwen3 Reranker 8B41K$0.135$0.00
AlibabaQwen3 Embedding 4B33K$0.027$0.00
AlibabaQwen3 Embedding 0.6B33K$0.0135$0.00