Qwen3.8-Max’s AI Numbers And The Complexity Of Ranking Performance

📊 Full opportunity report: Qwen3.8-Max’s AI Numbers And The Complexity Of Ranking Performance on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the broad availability of Qwen3.8-Max, confirming it has 2.4 trillion parameters with roughly 95 billion active per query. Its benchmark scores show strong performance, but the model’s ranking complexity remains nuanced and selective.

Alibaba has officially released the full benchmark table for its AI model Qwen3.8-Max, confirming it has 2.4 trillion parameters and approximately 95 billion active parameters per query. This marks a significant step in transparency and performance validation for one of the largest open-weight models to date, with implications for AI benchmarking and deployment strategies.

After weeks of speculation, Alibaba confirmed on August 3 that Qwen3.8-Max is a 2.4 trillion-parameter model built on the Qwen3.5 architecture, utilizing sparse mixture-of-experts technology. The model demonstrated its capabilities through a comprehensive benchmark table, which Alibaba published publicly for the first time, revealing scores across multiple AI evaluation suites.

The model’s active parameters are approximately 95 billion per query, representing about 4% of its total network, a key detail that clarifies its computational footprint. Benchmark results show that Qwen3.8-Max outperforms several competitors in specific tasks, such as Terminal-Bench 2.1 with a score of 86.6, surpassing Claude Fable 5 but trailing behind GPT-5.6 Sol at 88.8. It also leads in multimodal and agentic tasks, with top scores in OSWorld-Verified and Parametric CAD Bench.

Alibaba also disclosed that the model significantly improved in agentic tasks compared to its predecessor, DeepSWE, which jumped from 21.6 to 56.6 points, reflecting notable progress in long-horizon reasoning and environment interaction. However, on deep software-engineering benchmarks like SWE-bench Pro, the model remains behind Fable 5, with a score of 67.7 versus 80.0, indicating that the model’s strengths are selective and task-dependent.

At a glance
reportWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba officially released detailed benchmark scores and specifications for Qwen3.8-Max, confirming its 2.4 trillion parameters and high performance in key tests.
Crypto market snapshot
Fear & Greed Index
25/100 — Extreme Fear
Bitcoin BTC$63,655▲ 1.5%
Ethereum ETH$1,856▲ 0.4%
Tether USDT$0.9992▲ 0.0%
BNB BNB$590.02▲ 1.2%
USDC USDC$0.9996▲ 0.0%
XRP XRP$1.07▲ 0.5%
Solana SOL$73.4▲ 1.2%
TRON TRX$0.3287▲ 0.8%
Live data · CoinGecko · alternative.me (24h change)
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Benchmark Transparency

The release of detailed benchmark scores and the full spec sheet marks a milestone in AI transparency, allowing the industry to better understand the capabilities and limitations of one of the largest open-weight models. It underscores the ongoing challenge of comparing models with different architectures, parameter counts, and evaluation methods, especially when claims of performance are often selectively presented.

Furthermore, the distinction between the 2.4 trillion total parameters and the roughly 95 billion active parameters per query highlights the importance of understanding network efficiency and activation sparsity. This influences deployment strategies, especially for models intended for local or resource-constrained environments, such as the upcoming Qwen3.8-27B checkpoint designed for single-machine inference.

Overall, this development emphasizes that model size alone does not determine performance; the way parameters are utilized and evaluated remains critical, impacting future AI benchmarking and commercial deployment decisions.

Scaling AI: The AI Governance and Security Playbook for Executives

Scaling AI: The AI Governance and Security Playbook for Executives

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Releases

Alibaba’s AI model development has been characterized by strategic secrecy followed by targeted disclosures. In July, the company previewed Qwen3.8-Max during the World AI Conference in Shanghai, initially as an anonymous model called 'kaleb,' which was later identified as the stealth preview of Qwen3.8-Max. The company’s marketing emphasized its size—2.4 trillion parameters—and positioned it as 'second only to Fable 5,' though without providing detailed benchmarks at the time.

Prior to this, Alibaba’s models like Kimi K3 and earlier Qwen versions had gained attention for their scaling and multimodal capabilities. The recent benchmarks confirm that Alibaba’s approach involves sparse mixture-of-experts architectures, enabling large parameter counts while maintaining manageable active parameter levels during inference. The model’s performance in various benchmarks has been a focus of industry interest, especially given the lack of transparency around comparable models from other providers.

The current release follows a pattern of staged disclosures, with Alibaba gradually revealing more detailed specifications and performance metrics, culminating in the full benchmark table published on August 3.

"We are committed to transparency and advancing AI performance benchmarking to better serve developers and researchers."

— Alibaba spokesperson

Amazon

large language model performance monitor

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Deployment and Licensing

While Alibaba confirmed the benchmark scores and specifications, several details remain unconfirmed. The licensing terms for the 2.4 trillion-parameter weights are still unpublished, raising questions about open-source status and commercial restrictions. Additionally, the full performance of the 27B checkpoint, intended for local deployment, has not yet been publicly benchmarked, leaving uncertainty about its relative performance and agentic capabilities post-compression.

It is also unclear how the model’s performance will hold up in real-world applications outside controlled benchmarks, and whether the claimed improvements in agentic reasoning will translate into practical deployment scenarios.

Amazon

AI model parameter analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in Alibaba’s AI Model Strategy

Alibaba is expected to release the open weights for Qwen3.8-Max next week, along with further licensing details, which will influence how the model can be adopted by developers and organizations. The company may also publish additional benchmark results for the 27B variant, clarifying its suitability for local deployment.

Industry observers will be watching to see if the model’s agentic improvements are sustained in real-world tasks and whether Alibaba’s transparency prompts competitors to follow suit with more detailed disclosures. Further, the impact on AI benchmarking standards and the broader ecosystem’s understanding of large sparse models will be significant.

Amazon

AI benchmarking datasets

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the significance of Alibaba’s 2.4 trillion parameters?

The 2.4 trillion parameters indicate the model’s overall size, but only about 95 billion are active during inference, reflecting a sparse architecture that balances size with efficiency.

How does Qwen3.8-Max compare to other large models?

It outperforms some models like Claude Fable 5 in certain benchmarks but trails behind GPT-5.6 Sol at the top end. Its strengths are notable in multimodal and agentic tasks, though it remains selective.

Will the open weights be available for commercial use?

The weights are scheduled to ship next week, but licensing terms are still unpublished, so their commercial and open-source status remains uncertain.

What are the implications for local deployment?

The upcoming 27B checkpoint is designed for local inference on high-memory machines, but its performance relative to the flagship remains to be seen, especially regarding agentic capabilities after compression.

Why is benchmark transparency important?

It allows for fair comparison of models, helps identify strengths and weaknesses, and fosters trust in claims about AI performance, especially among developers and researchers.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

Inside The AI Data Future: OpenAI’s Enterprise Stack In 2026

OpenAI expands its enterprise offerings in 2026 with new data governance controls, emphasizing privacy, security, and integrated AI agents for business use.

Three Days at the Frontier: Washington Suspends Fable 5 and Mythos 5

The US government temporarily halts Anthropic’s Fable 5 and Mythos 5 models over security concerns following a jailbreak demonstration, raising industry and geopolitical questions.

Kashmir, Jammu And Kashmir, India Surges In Global Coverage

Kashmir, Jammu and Kashmir, India experiences a significant increase in international media mentions, highlighting growing global interest in the region.

Germany set to restrict its Freedom of Information Act

Germany is preparing to introduce restrictions to its Freedom of Information Act, raising concerns over transparency and public access to government data.