Last updated: July 20, 2026

Qwen 3.8 is coming, but the final model is not fully released yet. Alibaba’s Qwen team announced Qwen 3.8 on July 19, described it as a 2.4-trillion-parameter model, and said it would go open-weight “soon.” The model people can use today is a changing preview called qwen3.8-max-preview, available through Alibaba’s Token Plan, Qoder, and QoderWork.

That distinction matters. Qwen 3.8 already has a real endpoint, documented context and output limits, reasoning controls, vision input, and subscription access. But Alibaba has not yet published the final weights, a release date for those weights, a license, a full model card, a conventional Qwen 3.8 pay-as-you-go price, or a complete official benchmark table.

The practical verdict is therefore: Qwen3.8-Max-Preview is worth testing, but it is too early to treat it as a stable production replacement for Kimi K3, GPT-5.6 Sol, or Claude Fable 5.

Source check Sources checked

Verified July 20, 2026

The endpoint is explicitly a preview and can change. Promotions, quotas, supported tools, and model behavior can also change without a new model name, so record the date and configuration of every evaluation.

Qwen 3.8 at a glance

QuestionCurrent answer
Is Qwen 3.8 released?The hosted qwen3.8-max-preview is live; the final production model and downloadable weights are not yet available.
When was it announced?July 19, 2026.
How large is it?2.4T total parameters, according to Qwen. Active parameters and architecture details have not been disclosed.
What is the context window?983,616 tokens in Qwen Cloud’s official integration metadata.
What is the maximum output?131,072 tokens in the same official metadata.
Where can you use it?Token Plan, Qoder, and QoderWork; Qwen Cloud also documents connections to tools such as Codex, Qwen Code, Claude Code, Cursor, and OpenCode.
How is it priced?Credits-based subscriptions start at a promotional $6/month for individuals; no ordinary Qwen 3.8 per-million-token list price is published.
Are benchmarks available?Alibaba claims frontier-level performance, second only to Claude Fable 5, but has not published the supporting benchmark table or methodology.
Is it open source?Not yet. Qwen says open weights are coming soon, but has not announced a date, license, or exact checkpoint.
Qwen 3.8 preview release, parameter, context, output, access, pricing, benchmark, and open-weight facts
Qwen3.8-Max-Preview facts and unresolved questions checked July 20, 2026.

Do not confuse Qwen 3.8 with Qwen3-8B. Qwen3-8B is an older eight-billion-parameter model. Qwen3.8 is the new model-generation name, and its preview has 2.4 trillion total parameters.

Official Qwen launch graphic saying Qwen 3.8 is coming soon with 2.4T parameters and open weights
Qwen announced a 2.4T Qwen 3.8 model and promised open weights, while releasing Qwen3.8-Max-Preview for immediate testing. Credit: Qwen / Alibaba · · Source

What is Qwen3.8-Max-Preview?

qwen3.8-max-preview is the first publicly accessible version of Alibaba’s next Qwen flagship. It is not simply a renamed Qwen3.7-Max endpoint. Qwen’s official announcement gives it a new generation number, a reported 2.4T parameter count, and a future open-weight commitment.

The word preview sets the current expectations:

  • Qwen Cloud says the model will be continuously improved during the preview period.
  • The preview may later be taken offline or replaced by a production version.
  • Behavior can change without the reproducibility of a dated, downloadable checkpoint.
  • Prices are promotional and expressed through Credits rather than a normal token rate.
  • The final license, architecture, active parameter count, training details, and serving requirements are still unknown.

This makes the preview useful for exploration, comparative testing, and early workflow design. It makes it less suitable for a production migration that assumes a pinned model will behave the same next month.

Qwen 3.8 specs and capabilities

Qwen has not published a full model card, but its live documentation exposes several concrete specifications.

SpecificationQwen3.8-Max-Preview
Model IDqwen3.8-max-preview
Reported total parameters2.4 trillion
Active parametersNot disclosed
Context window983,616 tokens
Maximum output131,072 tokens
Input in the Codex integrationText and images
ReasoningAlways enabled
Reasoning levelslow, high, xhigh
Default reasoning levelxhigh
Default thinking temperature0.6; lower values are adjusted to 0.6
Parallel tool calls in Codex metadataNot supported
Built-in Token Plan toolsWeb search, code interpreter, web scraping, reverse image search, text-to-image search

The Codex integration guide is the strongest current source for the exact limits. It lists a 983,616-token context window, text and image input, and an effective context setting of 95% for Codex. Qwen Cloud’s OpenClaw metadata also lists a 131,072-token maximum output.

Those are integration specifications, not a substitute for a final model card. Still, they are more reliable than third-party pages rounding the context to 991K or 1M without showing where the number came from.

Always-on reasoning changes the cost and latency picture

Qwen3.8-Max-Preview cannot turn thinking off in the documented Token Plan integrations. You can select low, high, or xhigh, and the default is xhigh.

That matters because two requests with the same visible prompt and answer can consume different Credits depending on reasoning depth, accumulated context, cache reuse, and tool calls. A low subscription price does not guarantee a low cost per accepted task if the model produces long reasoning traces or needs repeated corrections.

Use low for easy, verifiable tasks, then test high or xhigh only when the extra reasoning improves the final result enough to justify the added latency and usage.

How to access Qwen 3.8

Alibaba currently names three first-party access paths:

  1. Qwen Cloud Token Plan. Subscribe to an Individual or Team plan, create the plan-specific API key, and choose the exact qwen3.8-max-preview model ID in a compatible tool.
  2. Qoder. Alibaba’s coding environment exposes the preview for repository and software-development workflows.
  3. QoderWork. Alibaba’s desktop assistant exposes the model for documents, files, data processing, and broader knowledge-work tasks.

Qwen Cloud documents Token Plan integrations with Qwen Code, Codex, Claude Code, Cursor, OpenCode, Cline, OpenClaw, and other agent tools. This is API-shaped access, but the subscription terms restrict it to interactive use inside compatible tools.

To evaluate the preview responsibly:

  1. Use the exact model ID qwen3.8-max-preview.
  2. Record the test date because the endpoint can change during preview.
  3. Record the reasoning level; the default is xhigh.
  4. Keep the same prompt, repository snapshot, tools, permissions, and time limit when comparing models.
  5. Measure accepted-result quality, latency, total tokens or Credits, retries, and factual corrections.
  6. Re-run the evaluation after Alibaba publishes a production version or downloadable checkpoint.

Qwen 3.8 pricing: subscription Credits, not a normal API rate

There is no official standard Qwen 3.8 input/output price per million tokens in the current Qwen Cloud pay-as-you-go catalog. The preview is sold through Token Plan subscriptions, where Credits depend on model choice, tokens, cache behavior, thinking, and tool use.

Token Plan Individual pricing

Individual planPromotional monthly price5-hour limit7-day limitRecommended concurrency
Lite$6700 Credits2,500 Credits1–2 agents
Standard$183,000 Credits10,000 Credits3–4 agents
Pro$6812,000 Credits40,000 Credits6–8 agents

The regular prices shown behind the current promotion are $8, $25, and $80 per month. During the preview, Qwen Cloud says Qwen 3.8 can consume Credits at as little as 10% of the standard rate. Calls made from 22:00 to 08:00 UTC+8 receive an additional temporary 80% reduction on top of that promotion.

These discounts are not permanent rate-card promises. Qwen Cloud reserves the right to change the promotion, and the model itself may be replaced after preview.

Individual plans also have two simultaneous sliding windows. Reaching either the five-hour or seven-day limit pauses access until older usage falls out of the relevant window. The plan does not automatically switch to pay-as-you-go billing when the quota is exhausted.

Token Plan Team pricing

Team seatCurrent monthly priceMonthly quotaIntended usage level
Standard$20 per seat25,000 CreditsLight AI use
Pro$75 per seat100,000 CreditsFrequent AI coding
Max$200 per seat250,000 CreditsHeavy coding use
Shared package$700 per package625,000 CreditsOverage shared across seats

The Team plan does not use the Individual plan’s five-hour and seven-day windows. It uses monthly seat quotas, and Qwen Cloud says conversation data in Team Edition is not used for model training. The terms still limit the service to interactive use in compatible tools rather than automated scripts or application backends.

Both Individual and Team Token Plans currently use the Singapore region with Global deployment. Qwen Cloud warns that prompts and outputs involve cross-border data transfer, which is a procurement and compliance consideration rather than a footnote.

Why third-party Qwen 3.8 prices are misleading right now

Some aggregators display a per-token Qwen 3.8 price. Those figures may describe an aggregator’s own promotional route or convert Credits using assumptions that Qwen Cloud does not publish as a fixed rate.

Do not place an aggregator number beside Kimi’s official $3/$15 API rate, GPT-5.6 Sol’s official $5/$30 rate, or Claude Fable 5’s official $10/$50 rate and call it an apples-to-apples comparison. Until Alibaba publishes a standard pay-as-you-go rate for the production model, the honest Qwen 3.8 cost unit is subscription price plus Credits consumed per successful task.

Qwen 3.8 benchmarks: what is proven and what is not

Alibaba says Qwen 3.8 is compatible with leading frontier models and second only to Claude Fable 5. That is a positioning claim, not a reproducible benchmark result.

As of July 20, Alibaba has not published:

  • the benchmark names behind that ranking;
  • Qwen 3.8’s scores;
  • competitor configurations;
  • prompts, harnesses, retry policies, or reasoning settings;
  • a technical report or model card;
  • a versioned checkpoint that independent evaluators can rerun.

Therefore, headlines saying Qwen 3.8 “beats GPT-5.6 Sol” go beyond the public evidence. Alibaba’s ranking may eventually be supported, but the data needed to audit it is not public yet.

The strongest early independent coding signal

One documented matched repository test compared Qwen3.8-Max-Preview with Kimi K3 on a demanding software-architecture task. Each model inspected the same frozen set of 269 files, worked with the same tool and time constraints, and had to produce a cited integration design, data contract, migration plan, tests, risks, and evidence ledger.

ResultQwen3.8-Max-PreviewKimi K3
Score after factual penalties80/10083/100
Main strengthCleaner system boundaries and stronger replay metadataMore complete revision, regeneration, and lifecycle state
Tool behavior44 tool calls; no failed tool calls53 tool calls; two denied compound commands, followed by recovery

Both models completed the long, tool-based analysis and reached the same core architecture decision. Qwen used fewer tool calls and produced stronger provenance fields; Kimi delivered a more complete lifecycle design.

This was one session per model on different provider routes, so it cannot isolate model quality from serving, caching, or harness effects. Both reports also needed factual correction.

Treat this as evidence that Qwen 3.8 can sustain serious repository work—not proof that it universally beats Kimi, Claude, or GPT.

Qwen 3.8 vs Kimi K3, GPT-5.6 Sol, and Claude Fable 5

CategoryQwen3.8-Max-PreviewKimi K3GPT-5.6 SolClaude Fable 5
Current statusChanging hosted previewHosted model and API; full weights scheduledProduction proprietary modelProduction proprietary model
Total parameters2.4T reported2.8T reportedNot disclosedNot disclosed
Context983,616 in integration metadata1M1.05M1M
Max output131,072Not the deciding public launch detail128K128K
Price structureToken Plan subscription Credits$3 input / $15 output per MTok; $0.30 cached input$5 / $30 per MTok at standard short context$10 / $50 per MTok
Public benchmark evidenceNo complete official table; one useful matched repository testBroad vendor and third-party early resultsBroad official and third-party resultsBroad official and third-party results
Open weightsPromised “soon”; no date or licensePromised for July 27, 2026NoNo
Best current reason to testLow-entry-cost preview, long context, coding, vision, and agent toolsBetter-documented open-weight contender with API pricingStable flagship in ChatGPT, Codex, and OpenAI APIStrong coding and long-horizon work in Claude’s ecosystem

Qwen 3.8 vs Kimi K3

Kimi K3 is easier to evaluate as a production candidate today because Moonshot has published a technical overview, conventional API rates, a 1M context figure, broad benchmark results, and a date for the planned weight release. Qwen 3.8 is cheaper to enter through a $6 promotional subscription, but Credits make task cost harder to compare and the endpoint is explicitly moving.

Qwen’s appeal is that a 2.4T open-weight release from Alibaba could become a major counterweight to Kimi. The early repository test suggests the preview is already competitive. Kimi still wins on documentation completeness and current decision confidence.

For the Moonshot launch and current capacity situation, read the Kimi K3 subscription and benchmark guide.

Qwen 3.8 vs GPT-5.6 Sol

GPT-5.6 Sol has stable model IDs, published prices, a system card, documented tools, and a large official benchmark table. Qwen’s vendor claim places Qwen 3.8 near or above leading frontier models, but the missing table prevents a credible score-by-score comparison.

Qwen is the more interesting experiment if future open weights, Alibaba’s tool ecosystem, or discounted interactive access matters. Sol is the safer first production evaluation if you need mature documentation, conventional API billing, and a pinned provider model today. See the GPT-5.6 Sol review for its current rates, tools, context limits, and benchmark evidence.

Qwen 3.8 vs Claude Fable 5

Alibaba itself calls Fable 5 the only model ahead of Qwen 3.8. Without the underlying evaluation, that claim should not decide a purchase.

Fable 5 is substantially more expensive at its published $10/$50 per-million-token rate, but it is a documented production model with established Claude Code integration and benchmark results. Qwen’s Token Plan can be far cheaper for interactive experiments, yet the preview may change and does not have a standard token price.

If the goal is a frontier model for a real repository, run both against the same issues and measure accepted patches, repair loops, review findings, latency, and cost. Model positioning is not a substitute for a repository-specific evaluation.

Will Qwen 3.8 be open source?

Qwen says Qwen 3.8 will go open-weight soon. “Open-weight” is the careful term because Alibaba has not promised the training data, full training code, or an OSI-approved software-style license.

Several essential questions remain unanswered:

  1. Which checkpoint will be released? The announcement does not say whether the downloadable model will be identical to qwen3.8-max-preview or a later production version.
  2. When will it arrive? “Soon” is not a release date.
  3. What license will apply? Usage limits can matter for commercial deployment and redistribution.
  4. What is the architecture? Total parameters alone do not reveal active parameters, expert routing, memory bandwidth, or inference cost.
  5. How will it be served? A 2.4T checkpoint is not a normal local-laptop model, even if quantized versions become available.

For scale only, raw weight storage for 2.4T parameters is roughly 4.8 TB at 16-bit, 2.4 TB at 8-bit, or 1.2 TB at 4-bit, before runtime overhead, caches, framework allocations, and redundancy. Sparse activation can reduce computation per token, but it does not make the full checkpoint fit on an ordinary consumer GPU.

The open-weight promise is still the most strategically important part of the announcement. A versioned checkpoint would let independent researchers audit behavior, providers compete on serving price, and teams reproduce tests instead of chasing a changing preview endpoint.

Who should test Qwen 3.8 now?

Decision point Should you test Qwen3.8-Max-Preview?
Best fit
  • You benchmark coding agents on real repositories
  • You want long-context text and image reasoning
  • You can tolerate a changing preview
  • You want a low-cost interactive evaluation
Use carefully
  • You need stable monthly capacity
  • You handle sensitive or regulated data
  • You depend on parallel tool calls
  • You need a model version to remain behaviorally fixed
Use another option
  • You need automated application-backend use through Token Plan
  • You require downloadable weights today
  • You need a published license and full model card
  • You require conventional per-token production pricing
Qwen 3.8 is currently an evaluation target, not a default production recommendation. Save inputs, outputs, tool logs, reasoning level, model ID, date, and accepted-result criteria so the test remains useful when the production model arrives.

The most useful early workloads are:

  • architecture reviews with evidence requirements;
  • multi-file codebase reading;
  • frontend implementation from screenshots;
  • document and spreadsheet analysis in QoderWork;
  • long-context research with web and scraping tools;
  • second-opinion reviews against Kimi, Claude, or GPT.

Avoid using a preview’s strongest demo as your only test. Include routine tasks, ambiguous tasks, failure recovery, factual verification, and cost tracking. A model that creates a beautiful first draft can still be expensive or unreliable when revisions begin.

Qwen 3.8 limitations to know before switching

  1. It is a moving preview. Qwen Cloud explicitly says capabilities will be updated and the endpoint may be removed or replaced.
  2. The benchmark ranking is not auditable. “Second only to Fable 5” has no public test table behind it yet.
  3. Token pricing is not transparent. Credits depend on tokens, context, cache, reasoning, and tools.
  4. Token Plan is not a general production API plan. Automated scripts, backends, batch jobs, and scheduled background tasks are prohibited.
  5. Open weights are still a promise. There is no file, date, license, or serving guide yet.
  6. Always-on thinking can raise latency and usage. Even easy tasks use reasoning, though effort is adjustable.
  7. Context capacity is not context quality. A 983,616-token window does not guarantee reliable retrieval across every position or task.
  8. The deployment path has compliance implications. Token Plan currently uses Singapore with Global inference and cross-border transfers.
  9. Tool behavior depends on the integration. Qwen’s Codex metadata marks parallel tool calls unsupported even though serial tool workflows are available.
  10. Total parameters do not equal active compute. Architecture and active parameter details are missing.

Final verdict: promising enough to test, incomplete enough to wait

Qwen 3.8 is a serious announcement. A 2.4T Alibaba model with nearly one million tokens of context, 128K output, image understanding, always-on reasoning, agent tools, and promised open weights could become one of the most important open-model releases of 2026.

But the release story is currently ahead of the evidence. The available model is qwen3.8-max-preview, not a final open checkpoint. Alibaba’s frontier ranking lacks published scores, the price is a promotional Credits system, and core architecture and licensing details remain absent.

The right move is not to dismiss Qwen 3.8 or declare it the new winner. Test the preview on work you can verify, compare it with Kimi K3, GPT-5.6 Sol, and Claude Fable 5 under matched conditions, and repeat the evaluation when Alibaba publishes the production model and weights.

For now, Qwen 3.8 wins on strategic potential and low-cost preview access. It does not yet win on documentation, reproducibility, or production certainty.

FAQ

Is Qwen 3.8 released?
Qwen3.8-Max-Preview is live through Alibaba’s Token Plan, Qoder, and QoderWork. The final production model and downloadable open weights are not yet released. Qwen Cloud says the preview will keep changing and may later be removed or replaced.
What is the Qwen 3.8 model ID?
The current exact model ID is qwen3.8-max-preview. Qwen Cloud warns users to match supported model IDs character for character rather than infer aliases or future version compatibility.
How many parameters does Qwen 3.8 have?
Alibaba reports 2.4 trillion total parameters. It has not disclosed active parameters, the full architecture, expert count, or routing details, so total parameter count should not be treated as a direct measure of speed or per-token compute.
What is the Qwen 3.8 context window?
Qwen Cloud’s official integration metadata lists a 983,616-token context window and 131,072-token maximum output for Qwen3.8-Max-Preview. The Codex metadata uses 95% as the effective context percentage, reserving headroom for operation.
How much does Qwen 3.8 cost?
Qwen 3.8 currently uses Credits-based Token Plan subscriptions rather than a published standard input/output token rate. Promotional Individual prices are $6, $18, and $68 per month for Lite, Standard, and Pro. Team seats are $20, $75, and $200 per month, with different quotas. Promotions and deduction rates can change.
Can I use Qwen 3.8 through an API?
Token Plan provides dedicated keys and OpenAI- or Anthropic-compatible endpoints for interactive coding and agent tools. However, the terms prohibit automated scripts, application backends, non-interactive batch processing, and scheduled production automation. It should not be described as unrestricted production API access.
Is Qwen 3.8 open source?
Not yet. Qwen says open weights are coming soon, but it has not published the weights, release date, license, exact checkpoint identity, or serving requirements. “Open-weight roadmap” is more accurate than calling the model open source today.
Does Qwen 3.8 beat GPT-5.6 Sol or Claude Fable 5?
There is not enough public evidence to say. Alibaba positions Qwen 3.8 as second only to Claude Fable 5, which implies a strong comparison with GPT-5.6 Sol, but the company has not released the supporting benchmark names, scores, configurations, or methodology. Run matched tests on your own work before choosing.
Is Qwen 3.8 the same as Qwen3-8B?
No. Qwen3-8B is an eight-billion-parameter member of the older Qwen3 model family. Qwen 3.8 is a newer generation name, and Alibaba reports 2.4 trillion total parameters for the Qwen3.8 model announced in July 2026.