Is Qwen3.8-Max Truly The Second Best AI? The Numbers Might Surprise You

📊 Full opportunity report: Is Qwen3.8-Max Truly The Second Best AI? The Numbers Might Surprise You on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, a 2.4 trillion-parameter AI model. Benchmark results position it as the second-best overall, but with notable limitations. The model’s open weights and capabilities are now confirmed, raising questions about its real-world performance.

Alibaba has officially released the full specifications and benchmark results for Qwen3.8-Max, confirming its position as the second-best AI model based on recent performance metrics. This marks a significant milestone in AI development, with the company providing detailed data after weeks of speculation.

On August 3, Alibaba made Qwen3.8-Max broadly available, revealing that the model contains 2.4 trillion parameters and is built on the Qwen3.5 architecture. The model employs sparse mixture-of-experts techniques and supports multimodal inputs, including text, images, and video, with text output. The active parameters per query are approximately 95 billion, indicating a model roughly equivalent to a 95B compute model wearing a 2.4T parameter network.

Benchmark results, obtained using Alibaba’s own testing harness, show that Qwen3.8-Max scores 86.6 on Terminal-Bench 2.1, surpassing models like Claude Fable 5 and Claude Opus 4.8, but trailing behind GPT-5.6 Sol, which scored 88.8. It also leads in PaperBench at 93.0 and performs strongly in multimodal and agentic tasks, with notable improvements over previous versions, especially in long-horizon agentic tasks.

However, the model underperforms significantly on deep software engineering benchmarks such as SWE-bench Pro (67.7 vs. Fable 5’s 80.0) and FrontierSWE (73.5 vs. 88.8), showing that the claim of being second only to Fable 5 applies selectively. The company also announced an open-weight checkpoint of 2.4 trillion parameters, set to ship next week, alongside a smaller 27B model optimized for local deployment.

At a glance
updateWhen: announced August 3, 2023; benchmark res…
The developmentAlibaba officially released detailed specifications and benchmark results for Qwen3.8-Max, confirming its position as the second-best AI model based on recent tests.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Results and Open Release

The official data confirms that Qwen3.8-Max is among the top-tier AI models in terms of raw benchmark scores, which could influence market positioning and developer adoption. The open release of the weights allows broader access and experimentation, potentially accelerating AI research and deployment. However, the model’s limitations in software engineering tasks highlight that its practical utility may vary depending on application focus.

This development matters because it sets a new benchmark for open-weight models and demonstrates Alibaba’s capacity to produce competitive AI at scale, impacting industry dynamics and competitive strategies.

Amazon

AI model benchmark comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Recent AI Model Launches

Over the past two weeks, Alibaba’s AI developments have been shrouded in secrecy, with the company previewing Qwen3.8-Max in stealth and only revealing key details at the World AI Conference in Shanghai on July 19. Prior to this, models like Moonshot’s Kimi K3 and the anonymous “kaleb” had stirred market interest. Alibaba’s initial claim of Qwen3.8-Max being “second only to Fable 5” was based on a limited preview without full benchmark data, leading to speculation about its true capabilities.

The recent release of detailed benchmark scores and specifications clarifies the model's performance and positions it within the current AI landscape, where GPT-5.6 remains the top performer on some metrics. The open weights and upcoming smaller models follow a pattern of Alibaba focusing on both high-end performance and accessible deployment options.

"Qwen3.8-Max demonstrates our commitment to advancing AI capabilities and transparency with comprehensive benchmark results."

— Alibaba spokesperson

Unconfirmed Aspects of Qwen3.8-Max’s Performance and Licensing

Details about the model’s licensing terms remain unpublished, raising questions about usage rights and commercial deployment. Additionally, the performance on software engineering benchmarks, where the model trails significantly, suggests that its practical utility in certain domains may be limited. Whether the agentic improvements will persist after quantization or in real-world scenarios is also still unclear.

Further testing and community feedback are needed to fully understand the model’s capabilities and limitations.

Next Steps for Alibaba and the AI Community

Alibaba will release the 2.4 trillion-parameter open weights next week, enabling researchers and developers to evaluate and deploy the model independently. The smaller 27B checkpoint will also become available, targeting local deployment scenarios. Industry observers will closely monitor how the model performs across diverse tasks and whether its agentic capabilities sustain in practical applications. Ongoing benchmarking and community testing will shape its adoption and competitive positioning.

Key Questions

What makes Qwen3.8-Max stand out compared to other models?

It features 2.4 trillion parameters, supports multimodal inputs, and has demonstrated top performance on several benchmark tests, especially in agentic and multimodal tasks, according to Alibaba’s published data.

When will the open weights be available for download?

Alibaba has announced that the open weights of Qwen3.8-Max will ship next week, enabling broader access for research and deployment.

How does Qwen3.8-Max compare to GPT-5.6?

Benchmark scores show GPT-5.6 still leads on some metrics, such as Terminal-Bench 2.1 with 88.8, but Qwen3.8-Max outperforms many models in other areas, making it a strong contender in the AI landscape.

What are the limitations of Qwen3.8-Max?

The model underperforms significantly on software engineering benchmarks and its licensing and practical deployment details remain unclear, which could limit its immediate utility in certain domains.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Is The AI Market Cooling? Prices Drop Due To Economic Struggles, Not Progress

Recent declines in AI-related memory prices are driven by consumer demand exhaustion, not supply recovery, signaling a market slowdown amid economic challenges.

How AI Will Impact Daily Life In 2026: Top 6 Insights

Exploring how artificial intelligence will reshape everyday activities by 2026, with six key insights on technology, work, and society.

RHEO: Paint With Light

RHEO is a new app that transforms touch into flowing, beautiful light displays, running on iPhone, iPad, and Apple Vision Pro, emphasizing calm and accessibility.

9 Best Mobile Workstation Laptops for Professional Workflows in 2026

Explore the nine best mobile workstation laptops for professional workflows in 2026, including specs, features, and ideal use cases.