Hacker News ★ 86 3 min

Qwen3.8 Max now ranked as the best overall model by agentic index

🔗 https://artificialanalysis.ai/?intelligence=agentic-index

📌 【Artificial Analysis 最新評測】Qwen3.8 Max 登頂 Agentic Index,成為目前最強代理模型

TL;DR:在 Artificial Analysis 的 Agentic Index 排名中,Qwen3.8 Max 已被評為綜合表現最佳的代理模型。

隨著 AI Agent(代理)技術從單純的對話轉向複雜的任務執行,如何衡量模型在「代理能力」上的表現成為業界關注焦點。根據 Artificial Analysis 最新的 Intelligence Index v4.1.1 數據顯示,Qwen3.8 Max 在代理任務的綜合評估中取得了領先地位。

🤔 評估指標的演進與新標準

Artificial Analysis 透過其 Intelligence Index 對模型進行多維度評估,其中針對代理能力的評估(Agentic Index)是衡量模型在執行複雜任務時的智慧程度。目前的評估框架包含 9 項核心指標,用於量化模型的綜合實力:

  • GDPval-AA v2
  • 𝜏³-Banking
  • Terminal-Bench v2.1
  • SciCode
  • Humanity’s Last Exam
  • GPQA Diamond
  • CritPt
  • AA-Omniscience
  • AA-LCR

📊 Qwen3.8 Max 奪得代理能力桂冠

在對 595 個模型的廣泛評測中,Qwen3.8 Max 在 Agentic Index 中展現了卓越的綜合表現,被列為該指標下的最佳模型。這意味著在處理需要多步驟推理、工具調用或複雜邏輯的代理任務時,Qwen 系列模型具備極高的可靠性。

💡 模型開發趨勢觀察

從 Artificial Analysis 的數據分布可以看出,目前的 AI 競賽正從單純的「智慧度(Intelligence)」轉向「執行效率(Cost per Task)」與「代理能力(Agentic Capability)」的平衡。

  • 推理模型(Reasoning models):在評測中被特別標註,代表其具備更強的邏輯推演能力。
  • 成本效益:除了智慧度,工程師在選擇模型時,亦需考量「每項任務的平均成本(Weighted average cost per task)」,這包含輸入、快取(Cache hit/write)以及推理與回答的 token 費用。

🎯 實務啟示

對於開發 AI Agent 應用程式的工程師而言,Qwen3.8 Max 的表現提供了一個關鍵參考:在設計需要高度自主性與複雜任務處理能力的代理系統時,該模型目前在評測基準上具備極強的競爭力。

🔗 來源

#AI #Qwen #ArtificialAnalysis #AgenticAI #LLM #MachineLearning #AIBenchmarks #GenerativeAI #MachineLearningEngineering #AIModelRanking

原始資料 Hacker News · 收集於 2026-08-07
來源原標題
Qwen3.8 Max now ranked as the best overall model by agentic index
作者
apitman
原始標籤
hackernews
來源訊號
HN points 480 HN 留言 303
原始連結
https://artificialanalysis.ai/?intelligence=agentic-index

摘要原文

Artificial Analysis K Independent analysis of AI Understand the AI landscape to choose the best model and provider for your use case Update Intelligence Index v4.1.1 Intelligence Index v4.1.1 moves 𝜏³-Banking to v1.0.1 and upgrades the grader for HLE, AA-LCR, and AA-Omniscience to GPT-5.6 Luna (medium) Launch Endpoint Accuracy Index Measuring whether provider endpoints serve the same model quality as the reference Highlights Intelligence Artificial Analysis Intelligence Index · Higher is better Speed Output tokens per second · Higher is better Cost per Task Weighted average cost (USD) per Intelligence Index task · Lower is better Personalized model recommender Get personalized recommendations based on your priorities for intelligence, speed, and cost Explore agents for general work, coding, customer support, and more Compare AI agents across capabilities, pricing, and platform support Explore premium plans Access expanded benchmark data, custom visualizations, industry reports, and more Changelog New article published · 6 Aug Launching v4.1.1 of the Artificial Analysis Intelligence Index Methodology updated · 6 Aug Artificial Analysis Intelligence Index v4.1.1 New language model evaluation · 6 Aug Ling 3.0 Tiny New article published · 5 Aug Muse Spark 1.2 New language model evaluation · 5 Aug Qwen3.8 Max New language model evaluation · 5 Aug Ling-3.0-flash New language model evaluation · 5 Aug Muse Spark 1.2 (xhigh) New article published · 4 Aug Launching the Endpoint Accuracy Index: Same Model, Different Accuracy New language model evaluation · 3 Aug G9v3-39A5B New article published · 31 Jul DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash New language model evaluation · 31 Jul Celeris-1 New language model evaluation · 31 Jul DeepSeek V4 Flash 0731 (Reasoning, Max Effort) New article published · 30 Jul Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters Methodology updated · 30 Jul We have updated our Cost per Task methodology, resulting in slight absolute increases in cost estimates but with minimal impact on relative positioning. New language model evaluation · 30 Jul Kimi K3 (low) New language model evaluation · 30 Jul Inkling Small New article published · 29 Jul Agnes AI releases Agnes 2.5 Pro Alpha New article published · 24 Jul Claude Opus 5: the new leader in agentic knowledge work New article published · 24 Jul Opus 5: Fable 5 level intelligence at a lower cost per task New language model evaluation · 24 Jul Claude Opus 5 (Adaptive Reasoning, Low Effort) See more Intelligence Intelligence of leading AI models based on our independent evaluations Artificial Analysis Intelligence Index Updated Agentic Index Updated Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 26 of 595 models NEW Add model from specific provider Estimate (independent evaluation forthcoming) Reasoning models are indicated by a lightbulb icon Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights / Proprietary Reasoning / Non-Reasoning Text Only / Multimodal Inputs By Country Artificial Analysis Intelligence Index by Open Weights / Proprietary Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 26 of 595 models NEW Add model from specific provider Estimate (independent evaluation forthcoming) Proprietary Open Weights (Commercial Use Restricted) Open Weights Reasoning models are indicated by a lightbulb icon Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license). Cost per Task Time per Task Output Tokens per Task Cost per Intelligence Index Task Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better 26 of 595 models NEW Answer Reasoning Cache Write Cache Hit Input Reasoning models are indicated by a lightbulb icon Cost per Intelligence Index Task Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. Intelligence Index vs. Cost per Task Intelligence Index vs. Time per Task Intelligence Index vs. Output Tokens per Task Intelligence Index vs. Cost per Intelligence Index Task Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task 26 of 595 models NEW Most attractive quadrant Pareto line Xiaomi Meta Google Anthropic MiniMax OpenAI NVIDIA Alibaba SpaceXAI Mistral Z AI Kimi DeepSeek Reasoning models are indicated by a lightbulb icon Cost per Intelligence Index Task Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Frontier Language Model Intelligence, Over Time Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 14 of 57 model creators NEW Anthropic OpenAI Kimi Alibaba Meta SpaceXAI Z AI DeepSeek Google MiniMax Xiaomi Thinking Machines Mistral Cohere Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Coding Agent Index Performance, cost, and execution time for leading coding agents on end-to-end software engineering tasks Index Cost Execution Time Artificial Analysis Coding Agent Index Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better Color by Model Agent 15 of 52 models NEW Image & Video Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals Text to Image Image Editing Text to Video Image to Video Video Editing Text to Image Leaderboard Elo scores from blind preference votes in our Image Arena. See the full leaderboard here. (480 points, 303 comments on Hacker News)

tencent/hy3:free 自動生成