[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
https://www.latent.space/p/ainews-qwen-38-max24t-and-27b-new📌 【Alibaba Qwen 重磅發布】2.4T 參數巨獸 Qwen 3.8 Max 登場,開源界的新天花板?
TL;DR:Alibaba 發布 2.4T 參數旗艦模型 Qwen 3.8 Max,主攻長程代理人任務與多模態推理,並承諾下週釋出開源權重。
在經歷了去年的轉型與管理層變動後,市場曾對 Qwen 能否持續提供具競爭力的模型持有疑慮。隨著 Qwen 3.8 Max 的問世,這份疑慮已蕩然無存。這不僅是一個參數規模達 2.4T 的巨型模型,更展示了 AI 在長程任務(Long-horizon)與自主協作上的驚人潛力。
🧩 2.4T 參數規模與 MoE 架構設計
根據第三方技術摘要,Qwen 3.8 Max 採用了大規模的混合專家模型(MoE)架構:
- 總參數規模:2.4T
- 每次 Token 啟用的參數(Active Parameters):95B
- 啟動比率:約為 4%
- 支援能力:具備原生多模態規劃能力,視覺資訊直接整合進執行迴圈,而非僅作為輸入通道。
📊 自主協作與長程任務的極限測試
Qwen 3.8 Max 的設計核心在於處理需要長時間、多步驟的複雜工作流,其展示的案例包括:
- 自主編程:具備超過 10 天的無人值守編程能力,並能建立自我演進的編程測試框架。
- 自主研究:在 125 小時的迭代研究迴圈中,自主發明了一種新的資料選擇方法,在基準測試中超越原論文結果 2.71 分。
- 晶片設計優化:執行完整的矽設計流程(從 RTL 編輯到實體佈局),在滿足 500 MHz 時序收斂的前提下,將閘門數量從 8,298 降至 678,並減少 81% 的晶圓面積。
- 電商策略執行:在為期 365 天的模擬營運中,透過博弈論談判與庫存規劃,實現了 4.16 倍的投資報酬率。
🚀 基準測試表現:直逼 Claude Opus 等頂尖模型
在多項關鍵指標上,Qwen 3.8 Max 展示了與西方頂尖閉源模型並駕齊驅的實力:
| 測試項目 | Qwen 3.8 Max 表現 | 對比參考 |
|---|---|---|
| Frontend Code Arena | 排名第 4 (1,668 Elo) | 僅次於 Claude Opus 5 [Max] 與 Kimi K3 |
| Vision Arena | 排名第 2 (1,305) | 僅落後 Claude Fable 5 [High] 13 分 |
| SWE-bench | 87.3% | 高於 GPT-5.5 (82.6%) 與 GLM-5.2 (83.3%) |
| Vals Index | 66.1 (開源模型第 2 名) | 與 Claude Opus 4.7 表現持平 |
⚠️ API 價格與可用性
目前該模型已透過 API 提供服務,價格極具競爭力:
- 輸入:$2.00 / M tokens
- 輸出:$6.00 / M tokens
- 快取(Cached):$0.25 / M tokens
Alibaba 同時也宣布,除了旗艦級的 3.8 Max,Qwen 3.8-27B 也將於下週釋出開源權重(Open Weights)。
🎯 實務啟示
對於工程師而言,Qwen 3.8 Max 的出現標誌著開源模型已進入「巨型稀疏模型(Giant Sparse Models)」時代。其強大的長程代理人(Agentic)能力與多模態視覺反饋,意味著未來開發者可以利用這類模型來處理更複雜的自動化工作流,例如自動化軟體工程、複雜硬體設計或長週期的商務決策,而不再侷限於單次的指令對話。
🔗 來源
- 標題:[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
- 作者/機構:Latent Space
- 連結:https://www.latent.space/p/ainews-qwen-38-max24t-and-27b-new
#AI #LLM #Qwen #Alibaba #OpenWeights #MachineLearning #AgenticAI #Multimodal #Coding #TechNews
原始資料 Latent Space · 收集於 2026-08-04
摘要原文
After the Qwen Exodus last year and new management took over launching more closed model APIs, there was some real doubt as to whether or not this leading open models lab would continue to release relevant models. That doubt is now gone. Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered . Qwen offers them on API for $2 input/$6 output per million tokens, but they have promised to open-weight both models. Autonomous Long-Horizon Coding: 10+ Days Unattended Coding: Built a self-evolving coding harness from scratch over a multi-week autonomous run. Autonomous AI Research: Rebuilt a complete paper’s pipeline ( Unified Data Selection for LLM Reasoning ) from scratch, then autonomously ran an iterative research loop over 125 hours to invent a new data selection method beating the original paper’s benchmark by +2.71 points . Competitive Data Science: Competed against 526 human teams in the WWW2025 Multimodal Dialogue Intent Recognition Challenge, placing in the top 13% ( outperforming 87% of human teams ) within 24 hours. Autonomous Hardware & Chip Design: Executed a complete silicon design flow (GCD/RSA cryptographic accelerator) from RTL editing to simulation, synthesis, and physical layout. Reduced gate count from 8,298 to 678 gates while achieving an 81% die area reduction and meeting physical timing closure at 500 MHz. Deep Real-World Work & Operations: Demonstrated production-grade outputs across hundreds of professional workflows (e.g., corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research). Outperformed competing models in the E-Commerce Bench (a 365-day store operation simulation), generating a 4.16x return (¥416,252 balance) through continuous game-theoretic negotiation and inventory planning. Multimodal Agents & Visual Feedback: Integrates native visual feedback across planning, coding, and GUI interaction, enabling direct application recreation across platforms (desktop, mobile, web). Released Qwen-MM-Plugins to extend multimodal capabilities to existing agent frameworks. A very nice win for open weights! On today’s pod with Baseten we talked about what it’s like to support these massive model drops on release. AI News for 7/25/2026-7/27/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! Top Story: Qwen 3.8 Max open model launch Alibaba Qwen announced Qwen3.8-Max as its new flagship and said open weights are coming next week. Alibaba introduced Qwen3.8-Max as its “most capable model to date,” describing it as a 2.4T-parameter model focused on coding, long-horizon agentic work, and multimodal reasoning, with the explicit claim that open weights will be released next week , alongside Qwen3.8-27B also going open-weight @Alibaba_Qwen The launch tweet also included API pricing: $2.00 / M input tokens , $6.00 / M output tokens , and $0.25 / M cached tokens @Alibaba_Qwen Alibaba framed the model around several headline capabilities: 10+ days of autonomous coding , 500+ turns of chip design optimization , 365 days of e-commerce strategy , and native multimodal intelligence where vision is part of the execution loop rather than just an input channel @Alibaba_Qwen The company simultaneously pushed availability across its own surfaces and partners: Qwen Studio , API , Command Code , and later Venice ; infra and app builders quickly confirmed support plans or integrations including Baseten , Hermes Agent , and Command Code @Alibaba_Qwen @Alibaba_Qwen @baseten @Teknium The announcement landed as part of a broader pattern: multiple observers described it as evidence that the Chinese open-weight frontier is now competing directly with top Western closed models , especially in coding, agentic workflows, and multimodal tasks @kimmonismus @matvelloso Vendor-reported model details and performance claims were unusually aggressive for an open-weight release. 2.4T total parameters @Alibaba_Qwen Long-horizon agentic/cowork focus @Alibaba_Qwen Autonomous coding over 10+ days with a public GitHub trace @Alibaba_Qwen 500+ turns for chip design optimization @Alibaba_Qwen 365 days of e-commerce strategy execution @Alibaba_Qwen Native multimodal planning loop rather than vision-only input @Alibaba_Qwen Third-party summary tweet from ZhihuFrontier added more claimed or reported technical details: 95B active parameters per token , implying an MoE activation ratio of roughly 4% API exposes low / medium / xhigh reasoning-effort modes Compatibility with OpenAI and Anthropic protocols Benchmark claims: PaperBench 93.0 , CoWorkBench 74.8 , WideSearch 81.9 @ZhihuFrontier Vals AI independently posted concrete eval/runtime settings: Tested at temperature 0.7 with default top-p / top-k @ValsAI These numbers matter because they place Qwen3.8-Max in the same deployment class as other giant sparse open models like Kimi K3 and GLM-5.2 , not the more practical 30B–70B local tier. The model immediately posted strong third-party results, especially in coding-adjacent, vision, and design-heavy arenas. Frontend Code Arena: Qwen3.8-Max debuted at #4 overall with 1,668 Elo , trailing only Claude Opus 5 [Max] at 1,705 and Kimi K3 [Max] at 1,676 , and roughly tied with Claude Opus 5 [High] at 1,669 @arena In Frontend Code Arena subslices, it ranked: #3 Brand & Marketing, Reference-based design, Gaming, Content Creation Tools Vision Arena: Qwen3.8-Max ranked #2 with 1,305 , only 13 points behind Claude Fable 5 [High] @arena Vals Index: Qwen3.8-Max ranked #2 among open-weight models , #10 overall out of 43 , with a score of 66.1 @ValsAI It matched Claude Opus 4.7 on the Index, 66.1 vs 66.1 At about 2.3x lower cost per test : $2.68 vs $6.17 @ValsAI Vals’ benchmark-specific numbers: SWE-bench: 87.3% , ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%) , but behind Claude Opus 4.8 (89.2%) Terminal-Bench 2.1: 67.4 , up from 61.0 for Qwen 3.7 Max @ValsAI Vals also highlighted the pace of progress: Gain of 8.6 points in ~2.5 months Price cut from $2.50/$7.50 to $2.00/$6.00 input/output @ValsAI There were also more anecdotal but technically relevant claims: One user visualized benchmark deltas and argued “Opus 4.8 is mostly subsumed by 3.8-Max” on the chart they reconstructed @deliprao Another claimed Qwen 3.8 surpassed Fable 5 on Terminal Bench and said Anthropic was now under visible pressure @kimmonismus A separate tweet called Qwen 3.8 Max the “best object detection VLM” across satellite, infrared, documents, technical drawings, sketches, crowded scenes, and small objects, though this was based on examples rather than a cited benchmark paper @skalskip92 Facts / directly attributable claims Alibaba announced Qwen3.8-Max and said open weights arrive next week ; Qwen3.8-27B will also go open-weight @Alibaba_Qwen Alibaba disclosed API pricing of $2 input / $6 output / $0.25 cached per million tokens @Alibaba_Qwen Arena reported #4 in Frontend Code Arena at 1,668 and #2 in Vision Arena at 1,305 @arena @arena Vals reported 66.1 on Vals Index , #2 among open-weight models , 87.3% SWE-bench , 67.4 Terminal-Bench 2.1 , 1M context , 128k output , and lower cost-per-test than Opus 4.7 @ValsAI @ValsAI @ValsAI ZhihuFrontier stated 95B active parameters and protocol compatibility; this appears to be a secondary summary rather than an original Alibaba spec sheet @ZhihuFrontier Opinions / extrapolations / rhetoric “China is no longer lagging behind but competing on equal footing” @kimmonismus “Open models are winning now” @JonathanRoss321 “Looks like Opus 4.
由 tencent/hy3:free 自動生成