Alibaba releases Qwen3.8-Max 2.4T-parameter model and Qwen3.8-27B with open weights, emphasizing long-horizon agentic workflows
Qwen3.8-Max is positioned as a flagship model for autonomous coding, chip design, and multimodal agentic tasks, with open weights slated for release next week alongside Qwen3.8-27B.
1 source · cross-referenced
- Alibaba announced Qwen3.8-Max, a 2.4 trillion parameter model, and Qwen3.8-27B, both set to release open weights next week.
- The models emphasize long-horizon autonomous coding, multimodal agentic workflows, and competitive performance in coding and vision benchmarks.
- Qwen3.8-Max is priced at $2.00 per million input tokens and $6.00 per million output tokens on its API.
- Third-party evaluations place Qwen3.8-Max near the top of coding and vision leaderboards, outperforming several proprietary models in specific tasks.
Alibaba announced Qwen3.8-Max, a 2.4 trillion parameter model, and Qwen3.8-27B, both slated to release open weights next week. The company described Qwen3.8-Max as its "most capable model to date," focusing on coding, long-horizon agentic work, and multimodal reasoning. Qwen3.8-27B is also planned for open-weight release.
The models were framed around several headline capabilities, including autonomous coding over 10+ days, 500+ turns of chip design optimization, 365 days of e-commerce strategy execution, and native multimodal intelligence where vision is integrated into the execution loop. Alibaba also announced API pricing for Qwen3.8-Max at $2.00 per million input tokens and $6.00 per million output tokens.
Vendor-reported details include a total of 2.4 trillion parameters for Qwen3.8-Max, with claims of 95 billion active parameters per token, implying a roughly 4% MoE activation ratio. The model is reported to support a 1 million token context window and offers low, medium, and xhigh reasoning-effort modes via its API. Compatibility with OpenAI and Anthropic protocols was also noted.
Third-party evaluations placed Qwen3.8-Max near the top of coding and vision leaderboards. On Frontend Code Arena, it debuted at #4 overall with 1,668 Elo, trailing only Claude Opus 5 [Max], Kimi K3 [Max], and roughly tied with Claude Opus 5 [High]. In Vision Arena, it ranked #2 with 1,305 Elo, only 13 points behind Claude Fable 5 [High]. The Vals Index ranked Qwen3.8-Max #2 among open-weight models and #10 overall out of 43, with a score of 66.1, matching Claude Opus 4.7 on the Index at about 2.3x lower cost per test.
Benchmark-specific results highlighted Qwen3.8-Max's performance in coding and agentic tasks. On SWE-bench, it achieved 87.3%, outperforming GPT-5.5 (82.6%) and GLM-5.2 (83.3%) but trailing Claude Opus 4.8 (89.2%). On Terminal-Bench 2.1, it scored 67.4, up from 61.0 for Qwen 3.7 Max. Independent analyses noted a gain of 8.6 points in about 2.5 months, alongside a price reduction from $2.50/$7.50 to $2.00/$6.00 for input/output tokens.
The announcement coincided with broader industry reactions, with observers describing it as evidence that the Chinese open-weight frontier is now competing directly with top Western closed models, particularly in coding, agentic workflows, and multimodal tasks. The models were made available across Alibaba's surfaces and partners, including Qwen Studio, API, Command Code, and later Venice, with infra and app builders confirming support plans or integrations.
- Aug 4, 2026 · Simon Willison — everything
Steve Yegge describes how an AI coding agent’s ‘just two more things’ tic derailed a project
Trust79 - Aug 3, 2026 · Microsoft Research
Microsoft Research releases Orchard, an open framework for training and evaluating agentic AI systems
Trust84 - Aug 3, 2026 · AWS — Machine Learning Blog
Formula 1 reduces data source onboarding time from weeks to minutes using agentic AI on AWS
Trust79