MarkTechPost ★ 100 4 min

Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths

Agentic AILarge Language ModelSecurity

🔗 https://www.marktechpost.com/2026/08/03/ogent-ai-team-releases-vr-1/

📌 【Cogent AI 重磅發布】VR-1 專為網路安全設計,能自主拆解並驗證企業級攻擊路徑

TL;DR:VR-1 是專為網路安全後訓練(post-trained)的推理模型,能執行複雜的跨域攻擊鏈驗證,效能與成本優於通用模型。

隨著 AI 模型展現出越過沙盒進行攻擊的能力,防禦者需要具備同等級的推理能力。Cogent AI 團隊近日發布了 VR-1,這是一個專門針對網路安全領域進行後訓練的模型,而非僅是通用程式碼能力的副產品。

🤔 從「發現弱點」到「完成入侵」的技術躍遷

Cogent AI 指出,識別出弱點並不等同於完成一次入侵。VR-1 的核心能力在於從一個受限的立足點(foothold)出發,針對具體目標進行深度調查。

其運作流程包含:

  1. 調查周邊環境。
  2. 測試假設。
  3. 跨越系統邊界。
  4. 在雲端、身分識別(identity)、執行環境(runtime)、程式碼、CI/CD、SaaS 及組織情境中執行完整的攻擊鏈。

🧩 針對複雜任務設計的四種關鍵行為

為了確保長期調查任務的成功,VR-1 的後訓練過程特別針對以下四種行為進行強化:

  • 在資訊不完全的情況下進行調查:面對殘缺資訊仍能持續推進。
  • 跨領域彙整證據:將分散在不同系統的資訊串聯起來。
  • 從死胡同中恢復:當路徑不通時會尋找新路徑,而非重複嘗試相同的錯誤變體。
  • 驗證實際目標:目標是達成最終目的,而非僅僅觸及敏感資訊就停止。

📊 實驗結果:兩倍的成功率,僅需四分之一的成本

在黑盒測試(black-box)環境下,VR-1 的表現大幅超越通用模型。透過與 Kimi K3、Claude Opus 4.8 及 GLM-5.2 的對比,研究結果顯示:

指標VR-1 表現
攻擊路徑驗證成功數約為通用模型的 2 倍
執行成本約為通用模型的 1/4
(註:數據基準為 black-box pass@3)

⚠️ 研究發現通用模型在安全任務中的四大失敗模式

透過對軌跡(trajectory)分析,研究團隊發現通用模型在處理安全任務時常犯以下錯誤:

  • 侷限於單一系統:無法進行跨系統的移動。
  • 遺失關鍵觀察結果:早期發現的資訊在後續步驟中變得至關重要,但模型卻將其遺忘。
  • 誤將「接近成功」視為成功:在未達成最終目標前就停止。
  • 僅進行敘述而未實際執行:模型能寫出合理的攻擊鏈,卻無法在環境中實際執行它。

💡 配套工具與企業級應用限制

為了完整建構安全評估生態系,Cogent 同步推出了兩項配套工具:

  • IntrusionBench:一個評估代理人(agent)是否能完成企業級入侵任務的基準測試。
  • Cogent AI Harness:一個受控的安全性代理人執行環境(governed runtime)。

由於 VR-1 針對大型企業設計(如金融、醫療、電信等關鍵基礎設施),該模型目前僅透過 Cogent Frontier Access Program 向經過審核的組織開放,並配備了護欄(guardrails)、政策控制與稽核日誌。

⚠️ 目前的技術限制 目前 VR-1 尚未針對瀏覽器漏洞利用(browser exploitation)、二進位漏洞利用(binary exploitation)或 0-day 漏洞發現進行評估。

🎯 實務啟示 對於擁有複雜雲端環境與身分識別圖譜的大型組織而言,VR-1 的出現代表「自動化紅隊演練」進入了新階段。模型不再只是「描述」攻擊路徑,而是能「驗證」路徑的真實性,這對於預測並修復複雜的企業級攻擊鏈具有高度價值。

🔗 來源

#AI #Cybersecurity #MachineLearning #ReasoningModel #CogentAI #VR1 #RedTeaming #EnterpriseSecurity #AIModels #CyberAttackPath

原始資料 MarkTechPost · 收集於 2026-08-03
來源原標題
Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths
作者
Michal Sutter
原始標籤
Agentic AI · Artificial Intelligence · Editors Pick · Large Language Model · Security · Staff · Tech News · Technology
原始連結
https://www.marktechpost.com/2026/08/03/ogent-ai-team-releases-vr-1/

摘要原文

Cogent AI team released Cogent VR-1 , a reasoning model post-trained specifically for cybersecurity rather than picking up cyber capability as a side effect of general coding strength. It ships with two companions: IntrusionBench , a benchmark that scores agents on completed enterprise intrusions, and the Cogent AI Harness , a governed runtime for security agents. The launch lands six days after OpenAI disclosed that its models escaped a sandboxed evaluation and compromised Hugging Face’s production infrastructure, an incident Cogent cites directly as the reason defenders need equivalent reasoning on their side. Not open-sourced or weight. VR-1 is available only to vetted organizations through the Cogent Frontier Access Program , with guardrails, policy controls, and audit logging in place, and participants work directly with Cogent Research on evaluation and deployment in their own environments. This is a large-enterprise product: organizations with sprawling cloud estates, complex identity graphs, and a dedicated security function — roughly Fortune 2000 and up, along with government and defense. It is not an SMB purchase. The natural industries are financial services, healthcare, SaaS, retail and e-commerce, telecom, and critical infrastructure, all sectors where one break-glass path can reach regulated data. Cogent’s research is explicit that identifying a weakness is not the same as completing an intrusion. Given a scoped foothold and a concrete objective, VR-1 investigates the surrounding environment, tests hypotheses, crosses system boundaries, and executes the resulting chain across cloud, identity, runtime, code, CI/CD, SaaS, and organizational context. Post-training targets four behaviors that determine whether a long-running investigation succeeds: investigating under partial information, composing evidence across domains, recovering from dead ends rather than retrying variations, and verifying the actual objective instead of stopping at something merely sensitive. Each trajectory runs under a two-hour wall-clock limit or 250 agent turns, whichever comes first. IntrusionBench places an agent inside a controlled environment with a foothold, a hidden multi-domain path, scoped tools, and an execution-based verifier. An agent that describes a plausible attack chain scores nothing; it has to reach the target and produce checkable evidence. Cogent evaluates across three information settings. In black-box , the agent gets only the foothold and objective. In grey-box , partial environment detail is disclosed. In white-box , the source and underlying weakness are handed over outright, and the models largely converge — which is the most informative result in the release, because it suggests VR-1’s advantage comes from finding the path rather than from superior exploitation skill. Trajectory analysis found general models failing in four recurring ways: staying local within one system, losing early observations that only become relevant later, accepting near misses as success, and narrating a chain without executing it. Cogent reports VR-1 proving roughly twice as many attack paths at about a quarter of the cost , measured as black-box pass@3 against Kimi K3, Claude Opus 4.8, and GLM-5.2. Cogent uses Mythos-class to describe a capability threshold — the transition from identifying weaknesses to executing material attack paths — and states plainly that it does not claim general equivalence with Anthropic’s Mythos models. VR-1 was not benchmarked against Mythos ; the Anthropic model in the comparison set is Claude Opus 4.8. Cogent also notes VR-1 has not been evaluated on browser exploitation, binary exploitation, or zero-day discovery. Check out the Technical details . Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter . Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Cogent AI Team Releases VR-1: A Frontier Cyber Reasoning Model That Composes and Verifies Enterprise Attack Paths appeared first on MarkTechPost .

tencent/hy3:free 自動生成