From weeks to minutes: How Formula 1® uses agentic AI on AWS to accelerate data operations
https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations/📌 【AWS 案例研究】F1 如何利用 Agentic AI 將資料整合週期從數週縮短至 40 分鐘
TL;DR:F1 利用 Amazon Bedrock AgentCore 打造 Data Accelerator,實現資料來源自動化入庫,效率提升 95% 以上。
在賽車運動的世界裡,決策速度必須與賽車的速度同步。對於擁有 8 億粉絲的 Formula 1 (F1) 來說,行銷技術(MarTech)平臺「Customer 360」是連結粉絲與商業策略的核心神經系統。然而,面對海量的數位互動數據,傳統的手動工程流程已成為制約業務發展的瓶頸。
🤔 面對數據爆炸:18 個月的開發積壓與手動工程瓶頸
F1 的數據來源極其多元,包含售票夥伴、串流媒體、贊助商數據、社群媒體及周邊商品系統。然而,這套系統曾面臨三大嚴峻挑戰:
- 手動入庫效率極低:每新增一個數據來源,工程師必須手動編寫 Schema 映射、建立數據管道(Pipeline)、配置數據品質檢查、定義 GDPR 分類並設定治理政策。這導致每個來源平均耗時 6 到 8 週,團隊甚至累積了 18 個月的待辦清單。
- 上游結構變動頻繁:供應商經常在未通知的情況下更改欄位名稱或調整 Payload 結構,這些變動往往發生在關鍵的賽事週末或行銷活動期間,導致管線失效。
- 可視性破碎:日誌(Logs)分散在 S3、Amazon Redshift、Airflow 與 DBT 等不同服務中,當指標出現問題時,工程師必須花費數小時手動追蹤數據譜系(Data Lineage)。
🧩 Data Accelerator:基於 Agentic AI 的自動化解決方案
為了應對挑戰,F1 與 AWS 合作開發了「Data Accelerator」,透過 Amazon Bedrock AgentCore 實現端到端的自動化操作。這不是單純的程式碼生成器,而是一套具備「模組化技能」的 Agent 架構。
該方案透過五個工作流同時運行,其中核心的 Agentic 工作流包含以下階段:
第一階段:從需求文件到自動化 Pull Request
- 團隊將包含數據來源資訊的商業需求文件 (BRD) 上傳至 Amazon S3。
- AWS Lambda 觸發 Amazon Bedrock AgentCore Runtime,Agent 讀取 BRD 並生成配置檔案。
- Agent 透過 GitHub App 將檔案提交為 Pull Request (PR),並透過 Jira API 自動建立對應的票券。
- 工程師僅需進行審核與調整。
第二階段:自動化基礎設施與轉換
- 一旦配置核准,Agent 會進一步生成三組獨立的 PR,分別對應基礎設施 (Infrastructure)、數據轉換 (DBT) 與治理 (Governance) 儲存庫。
💡 具備合規能力的智慧 Agent 與一般工具不同,此 Agent 整合了 GDPR 分類功能。它會主動分析每個數據欄位,判斷是否包含個人隱私或敏感資料,並直接將標籤發佈至 SageMaker Unified Studio 的治理註冊表,讓合規團隊能即時掌握狀況。
🧩 模組化技能與多輪推理架構
該系統採用模組化設計,每個 Agent 擁有一系列獨立技能,例如:
- Schema 映射與資料類型推論
- 數據品質驗證
- 治理執行
- 敏感資料分類
在執行時,Agent 會透過「多輪推理過程 (Multi-pass reasoning process)」來精確執行任務:
- Pass-0:處理 Token 管理與清理。
- Pass-1:總結工具輸出結果。
- Pass-2:彙整整體評估,逐步提升準確度與完整性。
📊 從數週縮短至 40 分鐘的效能飛躍
透過 Data Accelerator 的導入,F1 取得了顯著的營運成效:
| 指標 | 導入前 (Manual) | 導入後 (Agentic AI) |
|---|---|---|
| 新數據源入庫時間 | 6 至 8 週 | 約 40 分鐘 (程式碼生成) + 數小時 (部署與審核) |
| 自動化程度 | 人工手動處理 | 95% 以上由 Agent 自主處理 |
| 問題修復週期 | 以「天」為單位 | 以「小時」為單位 |
此外,該系統還具備「Schema 演進監控」能力。當上游數據結構變動時,Agent 能透過 Amazon EventBridge 偵測事件,評估對下游管線的影響,並自動生成修復程式碼與 Jira 票券,讓工程師只需進行最終審核。
🎯 實務啟示
對於處理大規模、高動態數據的工程團隊,F1 的案例展示了 Agentic AI 的實戰價值:它不只是輔助寫程式,更是在複雜的企業流程中(如合規、治理、跨系統追蹤)扮演了自動化協作的角色,將工程師從重複性的基礎建設工作中解放,轉向更高價值的任務。
🔗 來源
- 標題:From weeks to minutes: How Formula 1® uses agentic AI on AWS to accelerate data operations
- 作者/機構:Subhro Bose @ AWS ML
- 連結:https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations/
#AI #AgenticAI #AWS #MachineLearning #DataEngineering #Formula1 #DataOps #AmazonBedrock #DataGovernance #Automation
原始資料 AWS ML · 收集於 2026-08-04
摘要原文
Formula 1 ® (F1) engages an audience of over 800 million fans globally across digital platforms, F1 TV, social media, ticketing, and merchandise year-round. Races happen every two weeks. Fan engagement windows are measured in minutes and commercial decisions need to move at the speed of the grid. Behind the scenes, F1’s marketing technology (MarTech) platform, Customer 360, captures interactions across all of these touchpoints to power personalization, segmentation, and commercial strategy. However, the platform faced a significant operational challenge. According to Chris Roberts, Director of IT at Formula 1, “Our MarTech platform is the nervous system of F1’s fan engagement. But every new data source required 6 to 8 weeks of manual engineering. We had an 18-month backlog just to integrate 12 new sources.” The business was generating data faster than the engineering team could wire it up. As a result, Matt Kemp, F1 Head of Data Operations, set to improve efficiencies and data quality. “Manually ingesting data sources is time consuming, creates solution variances, and ultimately results in data integrity issues. I wanted a solution that was repeatable, robust and reliable. AWS worked backwards from our needs to implement an agentic solution that worked end to end, applying business logic at each step.” In early 2026, F1 and AWS worked together to build the Data Accelerator, a solution that uses agentic AI on Amazon Bedrock AgentCore to transform F1’s MarTech data platform from a manually maintained system into a self-managed, observable, and unified data estate. In this post, we show how the Data Accelerator reduced data source onboarding from up to 8 weeks to approximately 40 minutes of code generation plus hours of deployment. It also identified and fixed data source anomalies in production, tracked data platform operations and agent lineage in a single window, and opened a gateway for analysts, engineers, and scientists to collaborate. “For the first time, we have end-to-end visibility across the entire MarTech platform with data lineage and root cause analysis, not just dashboards full of alerts,” says Roberts. F1’s Customer 360 platform ingests data from ticketing partners, streaming integrations, sponsor activation feeds, social media, and merchandise systems. Operating a data estate of this breadth and velocity surfaced three areas of friction the team set out to solve. First, onboarding each new data source was a heavily manual effort: engineers wrote schema mappings, built ingestion pipelines, configured data quality checks, defined General Data Protection Regulation (GDPR) classifications, and set governance policies by hand. This process took 6 to 8 weeks per source. Second, the platform had to keep pace with constantly evolving upstream feeds. Providers frequently changed column names, added fields, or restructured and rescheduled payloads without notice. Those changes often surfaced at the worst possible moment, such as mid race-weekend or during a mission-critical campaign launch. Third, visibility was fragmented. Logs were scattered across services with no unified data lineage. When a stakeholder questioned a metric, engineers spent hours manually tracing the issue across Amazon Simple Storage Service (Amazon S3) paths, Amazon Redshift control tables, Airflow logs, and DBT outputs. The Data Accelerator addressed these challenges through five workstreams delivered simultaneously: A sixth workstream optimized the customer identity resolution algorithms that unify fan touchpoints across channels. The following sections describe each workstream in detail. The centerpiece of the Data Accelerator is a set of platform agents that take a Business Requirements Document (BRD) with limited information about the data source and produce a fully production-ready onboarding pipeline. This includes infrastructure code, data transformations, governance policies, and GDPR classification without a human writing a single line of boilerplate. The agents work in two phases: When a new data source needs onboarding, a team member uploads a BRD to an Amazon S3 bucket. The upload triggers an AWS Lambda function, which invokes Amazon Bedrock AgentCore Runtime, a capability of Amazon Bedrock AgentCore. The agent reads the BRD and generates a set of configuration files. It then accesses GitHub through a GitHub App to push these files as a pull request to the standardized Git repository, and accesses Jira through its REST API to create a ticket referencing the PR. All agent conversations and actions are traced in Amazon CloudWatch through built-in AgentCore observability. The assigned engineer reviews, adjusts if necessary, and approves. Phase 1 workflow: a BRD upload triggers the agent to generate config files and open a pull request Once the configuration files are approved, a human triggers the next stage. The agent takes the approved configuration and generates three separate Pull Requests: All three PRs link to a single Jira ticket for traceability. Engineers review each one across the Infrastructure, DBT, and Governance repositories and approve. Phase 2 workflow: the agent generates infrastructure, transformation, and governance pull requests What distinguishes this from a basic code generator is the integrated GDPR classification. The agent proactively analyzes every data column, determines whether it contains personal data, sensitive personal data, or pseudonymized data, and tags it with the appropriate GDPR category. These tags publish directly to the governance registry in SageMaker Unified Studio, giving the compliance team immediate visibility without manual review cycles. The system is not a tightly coupled agent graph. A single agent operates with modular skill definitions, each encapsulating a distinct capability: schema mapping and data type inference, data quality validation, governance enforcement, and sensitive data classification. At runtime, the agent evaluates incoming requirements and activates the relevant skills, composing them through a multi-pass reasoning process. Pass-0 handles token management through scrubbing, Pass-1 summarizes tool outputs, and Pass-2 rolls up an overall assessment, refining accuracy and completeness progressively rather than relying on a one-shot response. New capabilities ship as new skill modules without changing the core agent loop, keeping the architecture maintainable and composable as the platform grows. The result is onboarding time dropped from 6 to 8 weeks to approximately 40 minutes of code generation plus hours of deployment and review. AI agents now handle 95% of the work autonomously. Onboarding new data sources is one challenge, but keeping existing integrations healthy is another. Upstream providers frequently modify their data structures, from renaming a column to creating a new field. Previously, the F1 team discovered these changes when a pipeline failed, often during a live race weekend. The same agent architecture that handles onboarding now continuously monitors for upstream schema changes. When a provider modifies their data structure, the agent detects it through event-driven triggers using AWS Lambda and Amazon EventBridge. It assesses the downstream impact, identifying which pipelines are affected, and which consumers depend on the changed fields. It then generates the necessary code updates across all affected repositories and creates a Jira ticket with full context and linked PRs. Engineers receive a notification that explains what changed, describes the impact, and presents a proposed fix for review. End-to-end resolution now takes hours instead of days. Schema evolution agentic workflow Before the Data Accelerator, working with Customer 360 data required navigating multiple disconnected environments. Data engineers curated pipelines in one account. Data scientists who wanted to model fan behavior needed access to a separate account, and analysts operated in a third world entirely.
由 tencent/hy3:free 自動生成