Interconnects ★ 61 4 min

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier

🔗 https://www.interconnects.ai/p/latest-open-artifacts-23-laguna-s21

📌 【產業觀察】開源模型進入決定性時代:多個強大 MoE 模型湧現帕累托前緣

TL;DR:隨著訓練成本攀升,開發者正轉向大規模 MoE 架構,開源模型正挑戰效能邊界。

隨著模型訓練成本逐年以數個數量級增長,業界曾預測模型開發將走向整合。然而現況顯示,越來越多公司投入數億甚至數十億美元訓練強大的模型,並釋出開源權重。隨著 Token 需求持續攀升,開源模型在生態系中的角色正進入關鍵的決定性時代。

🧩 MoE 架構成為主流:從 Thinking Machines 到 Tencent

目前的趨勢顯示,混合專家模型(MoE)已成為處理大規模參數與高效率需求的核心架構:

  • Inkling (Thinking Machines):首款 975B-A41B 多模態 MoE 模型。它支援文字、圖像與音訊輸入,並輸出文字。雖然在同規模模型中未必是最強,但其設計目標是作為微調(fine-tuning)的強大基座,並透過其商業服務 Tinker 提供支援。此外還釋出了一個極具競爭力的 276B-A12B 版本。
  • Hy3 (Tencent):一個 295B-A21B 的 MoE 模型。效能較前代全面提升,且授權方式從限制性授權轉為 Apache 2.0。值得注意的是,該模型能透過專用工具與 Sol 作為裁判,證明一個擁有 50 年歷史的數學問題。

🚀 開源透明度與效能的競賽:Poolside 與 DeepSeek

  • Laguna S2.1 (Poolside):這款 118B-A8B 的 MoE 模型已連續三個月出現在技術追蹤中。它經過重新預訓練與後訓練,且模型大小適合運行於 DGX Spark。Poolside 採用了 OpenMDW 授權(類似 Apache 2.0 但對 AI 具備更完善的法律支持),並公開了完整的評估軌跡,展現了極高的透明度。此外還推出了 33B-A3B 的小版本 Laguna-XS-2.1。
  • DeepSeek-V4-Flash-0731 (DeepSeek):在 OpenAI 調降小型模型價格後,DeepSeek 立即更新了 V4 Flash 模型,在參數效能比(performance per parameter)上挑戰帕累托前緣(Pareto frontier)。

📊 中國技術力量與硬體在地化:Kimi 與 Meituan

  • Kimi K3 (MoonshotAI):這是近期最大的開源釋出之一。它採用非商業授權,要求推理與微調供應商必須簽署商業協議。這引發了關於政策工具是否會限制美國實體使用中國開源模型的討論。
  • LongCat-2.0 (Meituan):這是一款擁有 1.6T 參數的大型 MoE 模型。其技術亮點在於完全使用 Ascend 910 硬體進行訓練,成為首個完全在中國加速器上訓練的非玩具級(non-toy)模型。

💡 全球範圍內的技術佈局

  • Motif-3-Beta (Motif-Technologies):來自韓國,預覽版本為 314B-A13B MoE,引入了 GDLA 與 mHC 等架構創新。
  • Apertus-v1.5-70B (Swiss-AI):透過增加 2T Token 對 Apertus 1.0 進行持續預訓練。
  • Instella-MoE-16B-A3B-Think (AMD):AMD 利用其 Instinct 系列顯示卡訓練的 16B-A3B MoE 模型,並提供從 Base 到 SFT(監督式微調)、MidTrain 到 DPO(直接偏好優化)的所有階段檢查點(checkpoints)。

🎯 實務啟示

對於工程師而言,MoE 架構的普及意味著未來開發者可以利用較小的計算成本,獲得接近巨型模型的效能。同時,隨著像 Poolside 這樣具備高度透明度的開源模型出現,以及 AMD 等硬體商提供完整的訓練階段權重,開發者在選擇模型基座與進行微調時,將擁有更多元的技術路徑與硬體優化空間。

🔗 來源

#AI #MachineLearning #OpenSource #MoE #LLM #DeepLearning #ArtificialIntelligence #AIHardware #AIModels #TechTrends

原始資料 Interconnects · 收集於 2026-08-03
來源原標題
Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
作者
Florian Brand
原始連結
https://www.interconnects.ai/p/latest-open-artifacts-23-laguna-s21

摘要原文

Consolidation has been one of the paths that many astute observers predicted for the near-future of labs training models. It was labelled as inevitable, as training costs are increasing by orders of magnitude every year. Yet, as someone who in 2024 would’ve predicted consolidation really picking up come 2026 or 2027, where are we? We’re at a place where more companies are training strong models — easily investing hundreds of millions to billions of dollars in the total effort still — and an increasing number of organizations are releasing these models openly. The demand for tokens is incredibly high, and likely to increase as models get more efficient and unlock more possible use cases. All of these labs we thought would need to consolidate are realizing that building token machines is a likely path to value, and more companies will identify that source of value over time. The prime example is Thinking Machines — when they announced their company in February 2025, very few people would’ve put them in the bucket of an open models company, myself included. Now their open model finetuning service is making hundreds of millions in revenue per year and they’re releasing the best open-weight models built in the U.S.A. — ahead of the early leaders in NVIDIA with Nemotron and Arcee’s Trilogy. On the other side of the ecosystem is the sustained pace from the Chinese labs, with newer entrants like Xiaomi still accumulating mindshare in the broader AI economy. Having predicted consolidation for a long time, it now seems like a safer bet is to predict continued adoption, and try to imagine the role that open models play there. How much can revenue-share licenses like Kimi K3 stick? How much market share can open models take? We’re entering the decisive era. This is one of the most packed recaps of open models we’ve ever had, we’re excited! Inkling by thinkingmachines : The first model from Thinking Machines is a 975B-A41B multimodal MoE that supports text, images, and audio as inputs and produces text as output. While it is not the strongest model among peers (in China) in its size class, it is positioned to be a great base for fine-tuning, e.g., through their commercial offering, Tinker. They also release a smaller version (276B-A12B), which is really competitive for its size. Hy3 by tencent : A 295B-A21B MoE from Tencent. It improves over its predecessor across all metrics. Most notable, however, is the license change: While the previous version (covered in Artifacts 21 ) used a custom and rather restrictive license, Tencent switched to Apache 2 for this release. The model was also able to proof a 50 year old math problem (with a dedicated harness and Sol as a judge, although it is unclear how important the latter really is). Laguna-S-2.1 by poolside : Poolside quickly rose out of nowhere to become a frequent guest at Artifacts, marking its third appearance in three consecutive months. S2.1 is a newly pre- and post-trained version of the 118B-A8B MoE that fits on a DGX Spark, which brought it a lot of attention. Poolside also adopted the OpenMDW license, which is an Apache 2.0-like free license but has better legal backing for AI models specifically. The company also goes into more detail in its blog , which includes all the evaluation trajectories . This is a lot of transparency for an open model release! DeepSeek-V4-Flash-0731 by deepseek-ai : Just one day after OpenAI has dropped the prices of their smallest model by 80% , the whale dropped an update to their V4 Flash model, beating Luna at the pareto frontier. The bigger model is not updated yet, so it remains to be seen where it will land in terms of performance. For the initial V4 releases, the Flash version was the star of the show in terms of performance per parameter, while Pro was rather underwhelming. Kimi-K3 by moonshotai : This is the biggest open model release in some time, and we covered it in a separate post and a podcast episode . It was released under a noncommercial license, requiring inference and fine-tuning providers to enter into a commercial agreement. Kevin Xu and Graham Webster argue in a post that these licenses enable potential future government action against US entities doing business with Chinese AI companies: But if a US company needs a contract with Moonshot to provide the inference tokens that Kimi K3 generates, the picture looks different. Some of the policy tools US officials and others have debated as potential levers to restrict Chinese open model use would more clearly apply. LongCat-2.0 by meituan-longcat : The Chinese DoorDash is back again. This time, the company released another big MoE with 1.6T parameters. While the model itself is not the most capable for its size beyond benchmarks, it was trained entirely on Ascend 910s, making it the first non-Huawei, non-toy model trained entirely on Chinese accelerators. Other Chinese chips are mostly used for inference (if at all). Laguna-XS-2.1 by poolside : An update to the small (33B-A3B) MoE from Poolside. Motif-3-Beta by Motif-Technologies : A preview of a 314B-A13B MoE by the Korean Motif. This is by far the company’s most ambitious model, as it is considerably larger and introduces some architectural innovations like GDLA and mHC. Apertus-v1.5-70B by swiss-ai : A continued pre-train of the fully open-source Apertus 1.0, using 2T more tokens. Instella-MoE-16B-A3B-Think by amd : A 16B-A3B MoE trained by AMD on Instinct cards. AMD also provides all the different stages, from the base to the SFT checkpoints, as well as MidTrain and DPO .

tencent/hy3:free 自動生成