Last updated: July 20, 2026
Qwen 3.8 is coming, but the final model is not fully released yet. Alibaba’s Qwen team announced Qwen 3.8 on July 19, described it as a 2.4-trillion-parameter model, and said it would go open-weight “soon.” The model people can use today is a changing preview called qwen3.8-max-preview, available through Alibaba’s Token Plan, Qoder, and QoderWork.
That distinction matters. Qwen 3.8 already has a real endpoint, documented context and output limits, reasoning controls, vision input, and subscription access. But Alibaba has not yet published the final weights, a release date for those weights, a license, a full model card, a conventional Qwen 3.8 pay-as-you-go price, or a complete official benchmark table.
The practical verdict is therefore: Qwen3.8-Max-Preview is worth testing, but it is too early to treat it as a stable production replacement for Kimi K3, GPT-5.6 Sol, or Claude Fable 5.
Verified July 20, 2026
- Launch and positioning: Qwen’s official Qwen 3.8 announcement
- Current individual access, prices, quotas, and restrictions: Token Plan Individual
- Current team prices and data terms: Token Plan Team Edition
- Context, output, image input, and reasoning settings: Qwen Cloud’s Codex integration guide
- Built-in web, scraping, search, and code tools: Qwen Cloud Harness tools
- Early independent evidence: matched Qwen 3.8 vs Kimi K3 repository test
The endpoint is explicitly a preview and can change. Promotions, quotas, supported tools, and model behavior can also change without a new model name, so record the date and configuration of every evaluation.
Qwen 3.8 at a glance
| Question | Current answer |
|---|---|
| Is Qwen 3.8 released? | The hosted qwen3.8-max-preview is live; the final production model and downloadable weights are not yet available. |
| When was it announced? | July 19, 2026. |
| How large is it? | 2.4T total parameters, according to Qwen. Active parameters and architecture details have not been disclosed. |
| What is the context window? | 983,616 tokens in Qwen Cloud’s official integration metadata. |
| What is the maximum output? | 131,072 tokens in the same official metadata. |
| Where can you use it? | Token Plan, Qoder, and QoderWork; Qwen Cloud also documents connections to tools such as Codex, Qwen Code, Claude Code, Cursor, and OpenCode. |
| How is it priced? | Credits-based subscriptions start at a promotional $6/month for individuals; no ordinary Qwen 3.8 per-million-token list price is published. |
| Are benchmarks available? | Alibaba claims frontier-level performance, second only to Claude Fable 5, but has not published the supporting benchmark table or methodology. |
| Is it open source? | Not yet. Qwen says open weights are coming soon, but has not announced a date, license, or exact checkpoint. |

Do not confuse Qwen 3.8 with Qwen3-8B. Qwen3-8B is an older eight-billion-parameter model. Qwen3.8 is the new model-generation name, and its preview has 2.4 trillion total parameters.

What is Qwen3.8-Max-Preview?
qwen3.8-max-preview is the first publicly accessible version of Alibaba’s next Qwen flagship. It is not simply a renamed Qwen3.7-Max endpoint. Qwen’s official announcement gives it a new generation number, a reported 2.4T parameter count, and a future open-weight commitment.
The word preview sets the current expectations:
- Qwen Cloud says the model will be continuously improved during the preview period.
- The preview may later be taken offline or replaced by a production version.
- Behavior can change without the reproducibility of a dated, downloadable checkpoint.
- Prices are promotional and expressed through Credits rather than a normal token rate.
- The final license, architecture, active parameter count, training details, and serving requirements are still unknown.
This makes the preview useful for exploration, comparative testing, and early workflow design. It makes it less suitable for a production migration that assumes a pinned model will behave the same next month.
Qwen 3.8 specs and capabilities
Qwen has not published a full model card, but its live documentation exposes several concrete specifications.
| Specification | Qwen3.8-Max-Preview |
|---|---|
| Model ID | qwen3.8-max-preview |
| Reported total parameters | 2.4 trillion |
| Active parameters | Not disclosed |
| Context window | 983,616 tokens |
| Maximum output | 131,072 tokens |
| Input in the Codex integration | Text and images |
| Reasoning | Always enabled |
| Reasoning levels | low, high, xhigh |
| Default reasoning level | xhigh |
| Default thinking temperature | 0.6; lower values are adjusted to 0.6 |
| Parallel tool calls in Codex metadata | Not supported |
| Built-in Token Plan tools | Web search, code interpreter, web scraping, reverse image search, text-to-image search |
The Codex integration guide is the strongest current source for the exact limits. It lists a 983,616-token context window, text and image input, and an effective context setting of 95% for Codex. Qwen Cloud’s OpenClaw metadata also lists a 131,072-token maximum output.
Those are integration specifications, not a substitute for a final model card. Still, they are more reliable than third-party pages rounding the context to 991K or 1M without showing where the number came from.
Always-on reasoning changes the cost and latency picture
Qwen3.8-Max-Preview cannot turn thinking off in the documented Token Plan integrations. You can select low, high, or xhigh, and the default is xhigh.
That matters because two requests with the same visible prompt and answer can consume different Credits depending on reasoning depth, accumulated context, cache reuse, and tool calls. A low subscription price does not guarantee a low cost per accepted task if the model produces long reasoning traces or needs repeated corrections.
Use low for easy, verifiable tasks, then test high or xhigh only when the extra reasoning improves the final result enough to justify the added latency and usage.
How to access Qwen 3.8
Alibaba currently names three first-party access paths:
- Qwen Cloud Token Plan. Subscribe to an Individual or Team plan, create the plan-specific API key, and choose the exact
qwen3.8-max-previewmodel ID in a compatible tool. - Qoder. Alibaba’s coding environment exposes the preview for repository and software-development workflows.
- QoderWork. Alibaba’s desktop assistant exposes the model for documents, files, data processing, and broader knowledge-work tasks.
Qwen Cloud documents Token Plan integrations with Qwen Code, Codex, Claude Code, Cursor, OpenCode, Cline, OpenClaw, and other agent tools. This is API-shaped access, but the subscription terms restrict it to interactive use inside compatible tools.
To evaluate the preview responsibly:
- Use the exact model ID
qwen3.8-max-preview. - Record the test date because the endpoint can change during preview.
- Record the reasoning level; the default is
xhigh. - Keep the same prompt, repository snapshot, tools, permissions, and time limit when comparing models.
- Measure accepted-result quality, latency, total tokens or Credits, retries, and factual corrections.
- Re-run the evaluation after Alibaba publishes a production version or downloadable checkpoint.
Qwen 3.8 pricing: subscription Credits, not a normal API rate
There is no official standard Qwen 3.8 input/output price per million tokens in the current Qwen Cloud pay-as-you-go catalog. The preview is sold through Token Plan subscriptions, where Credits depend on model choice, tokens, cache behavior, thinking, and tool use.
Token Plan Individual pricing
| Individual plan | Promotional monthly price | 5-hour limit | 7-day limit | Recommended concurrency |
|---|---|---|---|---|
| Lite | $6 | 700 Credits | 2,500 Credits | 1–2 agents |
| Standard | $18 | 3,000 Credits | 10,000 Credits | 3–4 agents |
| Pro | $68 | 12,000 Credits | 40,000 Credits | 6–8 agents |
The regular prices shown behind the current promotion are $8, $25, and $80 per month. During the preview, Qwen Cloud says Qwen 3.8 can consume Credits at as little as 10% of the standard rate. Calls made from 22:00 to 08:00 UTC+8 receive an additional temporary 80% reduction on top of that promotion.
These discounts are not permanent rate-card promises. Qwen Cloud reserves the right to change the promotion, and the model itself may be replaced after preview.
Individual plans also have two simultaneous sliding windows. Reaching either the five-hour or seven-day limit pauses access until older usage falls out of the relevant window. The plan does not automatically switch to pay-as-you-go billing when the quota is exhausted.
Token Plan Team pricing
| Team seat | Current monthly price | Monthly quota | Intended usage level |
|---|---|---|---|
| Standard | $20 per seat | 25,000 Credits | Light AI use |
| Pro | $75 per seat | 100,000 Credits | Frequent AI coding |
| Max | $200 per seat | 250,000 Credits | Heavy coding use |
| Shared package | $700 per package | 625,000 Credits | Overage shared across seats |
The Team plan does not use the Individual plan’s five-hour and seven-day windows. It uses monthly seat quotas, and Qwen Cloud says conversation data in Team Edition is not used for model training. The terms still limit the service to interactive use in compatible tools rather than automated scripts or application backends.
Both Individual and Team Token Plans currently use the Singapore region with Global deployment. Qwen Cloud warns that prompts and outputs involve cross-border data transfer, which is a procurement and compliance consideration rather than a footnote.
Why third-party Qwen 3.8 prices are misleading right now
Some aggregators display a per-token Qwen 3.8 price. Those figures may describe an aggregator’s own promotional route or convert Credits using assumptions that Qwen Cloud does not publish as a fixed rate.
Do not place an aggregator number beside Kimi’s official $3/$15 API rate, GPT-5.6 Sol’s official $5/$30 rate, or Claude Fable 5’s official $10/$50 rate and call it an apples-to-apples comparison. Until Alibaba publishes a standard pay-as-you-go rate for the production model, the honest Qwen 3.8 cost unit is subscription price plus Credits consumed per successful task.
Qwen 3.8 benchmarks: what is proven and what is not
Alibaba says Qwen 3.8 is compatible with leading frontier models and second only to Claude Fable 5. That is a positioning claim, not a reproducible benchmark result.
As of July 20, Alibaba has not published:
- the benchmark names behind that ranking;
- Qwen 3.8’s scores;
- competitor configurations;
- prompts, harnesses, retry policies, or reasoning settings;
- a technical report or model card;
- a versioned checkpoint that independent evaluators can rerun.
Therefore, headlines saying Qwen 3.8 “beats GPT-5.6 Sol” go beyond the public evidence. Alibaba’s ranking may eventually be supported, but the data needed to audit it is not public yet.
The strongest early independent coding signal
One documented matched repository test compared Qwen3.8-Max-Preview with Kimi K3 on a demanding software-architecture task. Each model inspected the same frozen set of 269 files, worked with the same tool and time constraints, and had to produce a cited integration design, data contract, migration plan, tests, risks, and evidence ledger.
| Result | Qwen3.8-Max-Preview | Kimi K3 |
|---|---|---|
| Score after factual penalties | 80/100 | 83/100 |
| Main strength | Cleaner system boundaries and stronger replay metadata | More complete revision, regeneration, and lifecycle state |
| Tool behavior | 44 tool calls; no failed tool calls | 53 tool calls; two denied compound commands, followed by recovery |
Both models completed the long, tool-based analysis and reached the same core architecture decision. Qwen used fewer tool calls and produced stronger provenance fields; Kimi delivered a more complete lifecycle design.
This was one session per model on different provider routes, so it cannot isolate model quality from serving, caching, or harness effects. Both reports also needed factual correction.
Treat this as evidence that Qwen 3.8 can sustain serious repository work—not proof that it universally beats Kimi, Claude, or GPT.
Qwen 3.8 vs Kimi K3, GPT-5.6 Sol, and Claude Fable 5
| Category | Qwen3.8-Max-Preview | Kimi K3 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|---|
| Current status | Changing hosted preview | Hosted model and API; full weights scheduled | Production proprietary model | Production proprietary model |
| Total parameters | 2.4T reported | 2.8T reported | Not disclosed | Not disclosed |
| Context | 983,616 in integration metadata | 1M | 1.05M | 1M |
| Max output | 131,072 | Not the deciding public launch detail | 128K | 128K |
| Price structure | Token Plan subscription Credits | $3 input / $15 output per MTok; $0.30 cached input | $5 / $30 per MTok at standard short context | $10 / $50 per MTok |
| Public benchmark evidence | No complete official table; one useful matched repository test | Broad vendor and third-party early results | Broad official and third-party results | Broad official and third-party results |
| Open weights | Promised “soon”; no date or license | Promised for July 27, 2026 | No | No |
| Best current reason to test | Low-entry-cost preview, long context, coding, vision, and agent tools | Better-documented open-weight contender with API pricing | Stable flagship in ChatGPT, Codex, and OpenAI API | Strong coding and long-horizon work in Claude’s ecosystem |
Qwen 3.8 vs Kimi K3
Kimi K3 is easier to evaluate as a production candidate today because Moonshot has published a technical overview, conventional API rates, a 1M context figure, broad benchmark results, and a date for the planned weight release. Qwen 3.8 is cheaper to enter through a $6 promotional subscription, but Credits make task cost harder to compare and the endpoint is explicitly moving.
Qwen’s appeal is that a 2.4T open-weight release from Alibaba could become a major counterweight to Kimi. The early repository test suggests the preview is already competitive. Kimi still wins on documentation completeness and current decision confidence.
For the Moonshot launch and current capacity situation, read the Kimi K3 subscription and benchmark guide.
Qwen 3.8 vs GPT-5.6 Sol
GPT-5.6 Sol has stable model IDs, published prices, a system card, documented tools, and a large official benchmark table. Qwen’s vendor claim places Qwen 3.8 near or above leading frontier models, but the missing table prevents a credible score-by-score comparison.
Qwen is the more interesting experiment if future open weights, Alibaba’s tool ecosystem, or discounted interactive access matters. Sol is the safer first production evaluation if you need mature documentation, conventional API billing, and a pinned provider model today. See the GPT-5.6 Sol review for its current rates, tools, context limits, and benchmark evidence.
Qwen 3.8 vs Claude Fable 5
Alibaba itself calls Fable 5 the only model ahead of Qwen 3.8. Without the underlying evaluation, that claim should not decide a purchase.
Fable 5 is substantially more expensive at its published $10/$50 per-million-token rate, but it is a documented production model with established Claude Code integration and benchmark results. Qwen’s Token Plan can be far cheaper for interactive experiments, yet the preview may change and does not have a standard token price.
If the goal is a frontier model for a real repository, run both against the same issues and measure accepted patches, repair loops, review findings, latency, and cost. Model positioning is not a substitute for a repository-specific evaluation.
Will Qwen 3.8 be open source?
Qwen says Qwen 3.8 will go open-weight soon. “Open-weight” is the careful term because Alibaba has not promised the training data, full training code, or an OSI-approved software-style license.
Several essential questions remain unanswered:
- Which checkpoint will be released? The announcement does not say whether the downloadable model will be identical to
qwen3.8-max-previewor a later production version. - When will it arrive? “Soon” is not a release date.
- What license will apply? Usage limits can matter for commercial deployment and redistribution.
- What is the architecture? Total parameters alone do not reveal active parameters, expert routing, memory bandwidth, or inference cost.
- How will it be served? A 2.4T checkpoint is not a normal local-laptop model, even if quantized versions become available.
For scale only, raw weight storage for 2.4T parameters is roughly 4.8 TB at 16-bit, 2.4 TB at 8-bit, or 1.2 TB at 4-bit, before runtime overhead, caches, framework allocations, and redundancy. Sparse activation can reduce computation per token, but it does not make the full checkpoint fit on an ordinary consumer GPU.
The open-weight promise is still the most strategically important part of the announcement. A versioned checkpoint would let independent researchers audit behavior, providers compete on serving price, and teams reproduce tests instead of chasing a changing preview endpoint.
Who should test Qwen 3.8 now?
- You benchmark coding agents on real repositories
- You want long-context text and image reasoning
- You can tolerate a changing preview
- You want a low-cost interactive evaluation
- You need stable monthly capacity
- You handle sensitive or regulated data
- You depend on parallel tool calls
- You need a model version to remain behaviorally fixed
- You need automated application-backend use through Token Plan
- You require downloadable weights today
- You need a published license and full model card
- You require conventional per-token production pricing
The most useful early workloads are:
- architecture reviews with evidence requirements;
- multi-file codebase reading;
- frontend implementation from screenshots;
- document and spreadsheet analysis in QoderWork;
- long-context research with web and scraping tools;
- second-opinion reviews against Kimi, Claude, or GPT.
Avoid using a preview’s strongest demo as your only test. Include routine tasks, ambiguous tasks, failure recovery, factual verification, and cost tracking. A model that creates a beautiful first draft can still be expensive or unreliable when revisions begin.
Qwen 3.8 limitations to know before switching
- It is a moving preview. Qwen Cloud explicitly says capabilities will be updated and the endpoint may be removed or replaced.
- The benchmark ranking is not auditable. “Second only to Fable 5” has no public test table behind it yet.
- Token pricing is not transparent. Credits depend on tokens, context, cache, reasoning, and tools.
- Token Plan is not a general production API plan. Automated scripts, backends, batch jobs, and scheduled background tasks are prohibited.
- Open weights are still a promise. There is no file, date, license, or serving guide yet.
- Always-on thinking can raise latency and usage. Even easy tasks use reasoning, though effort is adjustable.
- Context capacity is not context quality. A 983,616-token window does not guarantee reliable retrieval across every position or task.
- The deployment path has compliance implications. Token Plan currently uses Singapore with Global inference and cross-border transfers.
- Tool behavior depends on the integration. Qwen’s Codex metadata marks parallel tool calls unsupported even though serial tool workflows are available.
- Total parameters do not equal active compute. Architecture and active parameter details are missing.
Final verdict: promising enough to test, incomplete enough to wait
Qwen 3.8 is a serious announcement. A 2.4T Alibaba model with nearly one million tokens of context, 128K output, image understanding, always-on reasoning, agent tools, and promised open weights could become one of the most important open-model releases of 2026.
But the release story is currently ahead of the evidence. The available model is qwen3.8-max-preview, not a final open checkpoint. Alibaba’s frontier ranking lacks published scores, the price is a promotional Credits system, and core architecture and licensing details remain absent.
The right move is not to dismiss Qwen 3.8 or declare it the new winner. Test the preview on work you can verify, compare it with Kimi K3, GPT-5.6 Sol, and Claude Fable 5 under matched conditions, and repeat the evaluation when Alibaba publishes the production model and weights.
For now, Qwen 3.8 wins on strategic potential and low-cost preview access. It does not yet win on documentation, reproducibility, or production certainty.
FAQ
Is Qwen 3.8 released?
What is the Qwen 3.8 model ID?
qwen3.8-max-preview. Qwen Cloud warns users to match supported model IDs character for character rather than infer aliases or future version compatibility.