Qwen3.8-Max Is Official: 95B Active, Open Weights Next Week — and Where It Actually Wins

Read time: ~7 minutes. Key facts:

  • Qwen3.8-Max officially launched August 3, 2026 — “the most capable model in the Qwen family to date.”
  • 2.4T total parameters, 95B active — the active count was undisclosed at the July preview; this is the number that decides real serving cost.
  • First Qwen-Max-class model to get open weights, “released next week.” Built on the Qwen 3.5 architectural foundation.
  • Available now via QwenCloud API.
  • Honest benchmark read: it tops PaperBench (93.0) and beats Opus 4.8 on FrontierSWE (73.5 vs 70.0) — but loses Terminal-Bench 2.1 to GPT-5.6 Sol (86.6 vs 88.8) and SWE-bench Pro to Fable 5 (67.7 vs 80.0).

Sourcing note: all figures are from Alibaba’s official Qwen3.8-Max announcement (published 2026/08/03), read directly from the page. Benchmarks are Alibaba’s own reported results — their stated test conditions are included below so you can weigh them properly. Links at the bottom.

When we covered the Qwen3.8-Max preview in July, the most useful thing we could say was what Alibaba hadn’t said: no active-parameter count, no benchmarks, no license, no weights date. That guide promised a follow-up when the real numbers landed. They’ve landed — and they’re more interesting than the headlines suggest.


1. The number that was missing: 95B active

From the official announcement:

Qwen3.8-Max — now available via QwenCloud: 2.4T parameters (95B active), with open weights releasing next week.

In July, “2.4T parameters” told you almost nothing about cost, because a sparse MoE only fires a fraction per token. 95B active is the number that matters. For scale: that’s roughly 2.3× Inkling’s 41B active and ~0.9× Kimi K3’s 104B — this is a heavyweight to serve, not a cheap-tier model.

It’s built on the Qwen 3.5 architectural foundation, scaled to 2.4T.


2. The bigger news: open weights for a Max model

This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week.

This is the genuinely significant part. Alibaba’s last several Max flagships shipped closed. Open-weighting a 2.4T Max-class model puts it in the same self-hostable category as Kimi K3 and Inkling — and at 95B active, serving it will be a datacenter proposition, but a possible one.

Caveat worth stating plainly: “next week” is the announcement’s wording, with no specific date, and no license terms were published in the post. Both matter before you plan around it. We’ll cover the local-run path once the weights and license are actually out and verifiable.


3. The full benchmark table (Alibaba’s own)

This is the table the July preview didn’t have:

BenchmarkQwen3.8-MaxOpus 4.8Fable 5GPT-5.6 Sol (max)Qwen3.7-Max
Terminal Bench 2.186.684.684.688.874.5
SWE-bench Pro67.769.280.064.660.6
DeepSWE 1.156.659.070.073.021.6
NL2Repo-Bench55.969.447.2
FrontierSWE73.570.088.840.7
MLS-Bench-Lite41.042.849.946.231.7
PaperBench93.080.388.890.564.8
AndroidBench75.169.884.574.056.5

Test conditions Alibaba states (important for weighing these): SWE-bench Pro was evaluated with the Claude Code harness, temp 1.0, top_p 0.95, 256K context, with “problematic tasks corrected and all baselines evaluated on the refined benchmark.” DeepSWE 1.1 used Claude Code and mini-SWE-agent harnesses, same sampling, 256K context, reporting the highest score among both.


4. The honest read

The HN headline going around is “best overall model by agentic index.” Alibaba’s own table says something more nuanced:

Where Qwen3.8-Max genuinely leads:

  • PaperBench 93.0 — the highest in the table, above GPT-5.6 Sol (90.5) and Fable 5 (88.8).
  • FrontierSWE 73.5 — beats Opus 4.8 (70.0).
  • Terminal Bench 2.1 86.6 — beats both Opus 4.8 and Fable 5 (84.6 each).

Where it doesn’t:

  • Terminal Bench: loses to GPT-5.6 Sol (88.8).
  • SWE-bench Pro: 67.7 vs Fable 5’s 80.0 — a 12-point gap, and it also trails Opus 4.8.
  • DeepSWE 1.1: 56.6, behind Opus (59.0), Fable 5 (70.0), and GPT-5.6 Sol (73.0).
  • NL2Repo / MLS-Bench-Lite: behind Opus and Fable 5 respectively.

The fair summary: Qwen3.8-Max is a legitimate frontier-tier model that wins on research-style and long-horizon agentic work (PaperBench, FrontierSWE) while trailing the leaders on hard SWE benchmarks. “Best overall” depends entirely on which index you weight — and an aggregate index can hide a 12-point SWE-bench gap.

The most striking column is actually the last one. Against its own predecessor Qwen3.7-Max, the jump is enormous: Terminal Bench 74.5 → 86.6, DeepSWE 21.6 → 56.6, FrontierSWE 40.7 → 73.5, PaperBench 64.8 → 93.0. Whatever the cross-vendor ranking, this is a very large generational step.

Standard caveat: these are Alibaba’s own reported numbers, and cross-lab agent benchmarks are harness-sensitive — note they ran SWE-bench Pro and DeepSWE through Claude Code harnesses and reported the best of two for DeepSWE.


5. What to do with this

  • Want to try it now? It’s live on the QwenCloud API. The preview guide covers the access mechanics (OpenAI- and Anthropic-compatible protocols).
  • Waiting to self-host? Watch for next week’s weights drop — and check the license the moment it publishes, since none was stated in the announcement.
  • Picking a coding model today? If your workload is hard SWE-style repo work, the table says Fable 5 and GPT-5.6 Sol still lead. If it’s long-horizon research/agentic work, Qwen3.8-Max’s PaperBench and FrontierSWE numbers are the best case for it.
  • Cost-sensitive? 95B active is heavy. For cheap agentic work, DeepSeek V4 Flash at $0.14/$0.28 is a very different price point.

The takeaway

Qwen3.8-Max’s official launch (August 3, 2026) fills in exactly what the July preview left blank: 2.4T total / 95B active, a full benchmark table, and — the real headline — the first open weights for a Qwen-Max-class model, due next week. Read the table honestly: it tops PaperBench (93.0) and beats Opus 4.8 on FrontierSWE, but trails GPT-5.6 Sol on Terminal-Bench and Fable 5 on SWE-bench Pro by 12 points. The generational leap over Qwen3.7-Max is the least ambiguous result in the whole table. Watch for the weights and the license next week — that’s when the self-hosting story starts.

For the access mechanics see our Qwen3.8-Max preview guide; for a cheaper agentic option, DeepSeek V4 Flash; for local Qwen coding today, Qwen 3.6 locally.

Sources

  • Qwen3.8-Max: A New Bar for Coding and Cowork — Qwen (2026/08/03) — official release; 2.4T parameters with 95B active; first open-sourcing of a Qwen-Max-class model with open weights releasing next week; built on the Qwen 3.5 architectural foundation; available via QwenCloud API; full benchmark table vs Opus 4.8 / Fable 5 / GPT-5.6 Sol (max) / Qwen3.7-Max (Terminal Bench 2.1 86.6, SWE-bench Pro 67.7, DeepSWE 1.1 56.6, NL2Repo-Bench 55.9, FrontierSWE 73.5, MLS-Bench-Lite 41.0, PaperBench 93.0, AndroidBench 75.1); test conditions (SWE-bench Pro via Claude Code harness at temp 1.0 / top_p 0.95 / 256K context with refined benchmark; DeepSWE 1.1 via Claude Code and mini-SWE-agent harnesses reporting the higher score)
  • Our July preview coverage — what was undisclosed at preview time (active params, benchmarks, license, weights date)
  • All benchmark figures are Alibaba’s own reported results under their stated harnesses; no license terms were published in the announcement and the weights date is stated only as “next week.” Verified August 7, 2026.