Improving Fable 5's biology safeguards
https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards📌 【Anthropic】優化 Fable 5 生物安全防護機制,大幅降低生物相關查詢的誤判率
TL;DR:透過改進安全分類器(Classifiers),Anthropic 將 Fable 5 的生物相關回退(fallbacks)減少了約 85%。
在開發具備頂尖能力的 Frontier Model 時,如何在「釋放科學潛力」與「防止生物武器風險」之間取得平衡,是 AI 產業最棘手的課題。Anthropic 針對其 Fable 5 模型進行了重大更新,旨在減少生物學領域的「誤判」,讓使用者在進行日常健康或教育諮詢時,不再頻繁被強制切換至效能較低的模型。
🤔 為什麼生物安全防護如此困難?
Fable 5 在某些複雜的生物任務上,表現甚至能超越專家。這種能力是一把雙面刃:
- 正面影響:能協助研究人員開發新型醫療方案。
- 潛在風險:惡意行為者可能利用其能力開發生物武器,例如透過合成生物學或基因編輯技術製造威脅。
由於生物領域存在高度的「雙重用途」(dual-use)特性——例如研究疫苗需要培養病原體、研發降血壓藥物(如 Captopril)需要研究蛇毒成分——這使得 AI 很難僅透過簡單的關鍵字來區分「科學研究」與「惡意意圖」。
🧩 從「全面阻斷」到「精準識別」的技術演進
為了降低風險,Anthropic 在 Fable 5 發布初期採取了極其保守的策略:對幾乎所有生物相關查詢進行阻斷,並將請求導向較低能力的 Opus 5 模型(即 Fallback 機制)。雖然這保護了安全性,卻造成了大量的「誤判」(False Positives),讓合法的教育或醫療諮詢也被阻斷。
為了解決這個問題,Anthropic 採取了以下步驟來優化安全分類器(Safety Classifiers):
- 重寫分類器憲法(Constitution):制定了一套規則集,幫助模型辨識受保護內容與允許內容的細微差異。
- 專家回饋與訓練:邀請內部與外部專家提供反饋,並根據新的憲法開發訓練資料集,重新訓練分類器。
- 調整分類邊界:
- 舊版狀態:分類邊界過於靠近左側(良性區),導致大量良性請求被納入「安全邊界」內而遭到阻斷。
- 新版狀態:分類邊界向右移動,模型能更精準地識別出哪些是真正的惡意或雙重用途內容,從而放寬對良性查詢的限制。
📊 實驗結果:生物相關回退減少 85%
透過這次對分類器與訓練資料的優化,Fable 5 在產品層面上的表現如下:
- 回退率降低:生物相關的「回退」(Fallback)現象減少了約 85%。
- 使用者體驗提升:使用者在進行解釋實驗室檢驗結果、理解症狀或學習生物學知識等日常教育任務時,將能獲得更完整的模型支援。
⚠️ 目前的限制與未來方向
儘管有了顯著進步,但針對「雙重用途」的風險仍未完全消除。目前 Fable 5 對於以下領域的請求,仍會回退至 Opus 5 以確保安全:
- 病毒學 (Virology)
- 毒理學 (Toxicology)
- 分子設計 (Molecular design)
這意味著 Fable 5 目前仍不適用於專業的生物研究或藥物開發。Anthropic 表示,未來將透過「信任的存取路徑」(trusted access pathways)來逐步縮小這個差距,為專業生物學家提供受控的尖端能力。
🎯 實務啟示
對於開發者與研究者而言,這項更新展示了「安全性」與「可用性」之間的權衡過程。當模型能力跨入生物、網路安全等高風險領域時,透過「分類器 + 降級模型(Fallback)」的架構,是目前產業在確保安全的前提下,最穩健的部署策略。
🔗 來源
- 標題:Improving Fable 5’s biology safeguards
- 機構/作者:Anthropic
- 連結:https://www.anthropic.com/news/improving-fable-5-s-biology-safeguards
#AI #MachineLearning #Anthropic #Fable5 #AI-Safety #Biology #BioTech #LLM #DualUse #AI-Governance
原始資料 Anthropic News · 收集於 2026-08-07
摘要原文
Product Announcements Improving Fable 5's biology safeguards Aug 7, 2026 We’re making updates to Claude Fable 5’s biology safeguards in a way that substantially reduces false positives. Fable 5 users will now experience many fewer “fallbacks”—where the system switches to a less capable model after they make a biology-related query. In our testing, this update reduced biology-related fallbacks by about 85% across our product surfaces. 1 Fable 5 will thus be able to assist with a wider range of biology tasks. In practice, users should see far fewer fallbacks on everyday health and educational questions—for example, interpreting lab results, understanding symptoms, and learning about biology in an educational context. Healthcare professionals will be able to receive more support from Fable 5 on clinical tasks. We believe the greatest opportunity for AI to positively affect the world is in biology and medicine, and we're investing significantly in building a responsible way to give biologists frontier access. Today, Fable still falls back to Opus 5 for requests we consider dual-use—including virology, toxicology, and molecular design—so it isn't yet usable for professional biology research and drug development. We're committed to closing that gap through trusted access pathways for frontier biology capabilities. Why we built strong biology safeguards Our objective is to get Fable 5’s frontier capabilities into the hands of as many of our users as possible, as quickly as possible. However, to do so, we need to manage the increasing risks that come with models this capable. One such risk is in the field of biology: Fable 5 can now outperform experts on some highly complex biological tasks and provide operational support on others. That means that it can provide genuine assistance to a researcher developing a new medical treatment (which is the reason we’re so keen to widen access to the model via both classifier improvements and trusted access programs). But in the wrong hands, those same capabilities could be used by a malicious actor, for example in developing a biological weapon. Our capability assessments show that Fable 5 could provide significant uplift to such an actor—that is, it could provide them with capabilities they could not find anywhere else. It’s often difficult to tell apart beneficial and harmful uses of AI in biology. For example, in some cases researching a treatment for a disease requires scientists to produce the dangerous compounds that cause that disease in the first place. This is most obvious for live vaccines, which require scientists to grow the same pathogen they’re aiming to prevent. It’s also the case for some medicines. To develop the drug captopril, which treats hypertension, scientists isolated toxic components of snake venom that crash blood pressure in humans. As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits. Sophisticated actors who wish to use our models to do harm know how to exploit this ambiguity to obscure their intent, making dangerous tasks look like ordinary research pursuits. The US Intelligence Community’s 2026 Annual Threat Assessment makes clear that such actors exist, and that advances in biotechnology including synthetic biology and genomic editing “could lead to novel biological threats.” It notes that several state actors likely maintain active offensive biological and chemical weapons programs—programs that could be accelerated by access to the raw capabilities of frontier AI models. Because of our concerns about these “dual-use” capabilities (those that could be used for beneficial or harmful purposes, and where the line between them is not always easy to draw), we intentionally launched Fable 5 with almost all biology queries blocked. This enabled us to make the model available for users in other domains. We knew this would be frustrating for legitimate biology users: it would result in a high number of false positives in the near term, where users asking biology-related questions would have their requests blocked and sent to a less capable model. Nevertheless, we chose to make this tradeoff because the cost of Fable being misused in a dual-use domain like biology could potentially be catastrophic . How our biology safeguards work One of the core ways we protect against misuse in biology is via safety classifiers: smaller, automated AI systems that detect when Fable 5 is asked to perform a safeguarded biology task, or produce a harmful output (we've previously written about our similar classifiers in the domain of cybersecurity). In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked. Developing precise, robust classifiers is not a straightforward task. For a classifier to work rapidly and consistently, it has to learn the difference between what we consider “in scope” and “out of scope” for the topics and queries we consider to be potentially harmful. It takes time and iteration to tune the classifiers, avoiding both false positives (where classifiers fire on out-of-scope content) and false negatives (where in-scope content is missed). We also require our classifiers to be robust to attempts to bypass them (known as jailbreaks), which requires even further research and testing. Starting with a very broad biology classifier meant that we could give our users access to Fable 5 while we continued our research aimed at refining it. The alternative—holding back the model until much more safeguards progress was made—would have delayed the model’s general access, and its potential benefits to our users, by weeks or months. Over the past several weeks, we've carefully rewritten the classifier’s constitution (which consists of a collection of rules to help the model discern between safeguarded and allowed content), taking care to carve out benign uses in detail. We solicited feedback on the changes from a diverse range of experts (both internal and external to Anthropic). We then developed updated training data for the classifier based on that constitution, and retrained it, and verified the new classifier would still generally trigger for harmful and dual-use research biology content but would now enable a wider range of benign and beneficial uses. As is illustrated in the diagram below, these updates meant that—compared to at the time of Fable 5’s launch—the classifier will trigger for many fewer benign biology-related requests. Illustration of our biology classifiers. Content that falls on the left-hand side of the classifier boundary is allowed; content that falls on the right-hand side is safeguarded (and is therefore blocked and sent instead to a less capable model);. Clearly harmful content (red), and content that is dual-use (orange), triggers the classifier and is blocked. We include a safety margin that includes content that is very likely benign but which is still blocked out of an abundance of caution (light green). Clearly benign content is in darker green. Upon its launch, Fable 5 had very broad classifiers (A) that triggered on a wide range of requests—even ones that were almost certainly benign (those in the safety margin). The classifier boundary is thus very far to the left-hand side of the diagram. The update we are announcing today (B) means that many more benign requests are allowed by the classifier, which has become better at discerning subtle differences between benign and dual-use queries. The classifier boundary in the diagram has therefore moved further to the right-hand side. Conclusions There’s still much more to be done to refine our safeguards.
由 tencent/hy3:free 自動生成