MarkTechPost ★ 94 4 min

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio

Audio Language ModelLanguage ModelLarge Language ModelMachine LearningVoice AI

🔗 https://www.marktechpost.com/2026/08/01/minimax-releases-minimax-h3-an-omni-modal-video-model-that-generates-15-second-2k-clips-with-native-stereo-audio/

📌 MiniMax H3 多模態 2K 影片生成

TL;DR:MiniMax H3 為統一多模態模型,支援 2K 15 秒立體聲影片,透過 API 提供即時服務。

🎣 當影片生成仍需分離文字、圖像、音訊的專家模型時,單一端點能否同時理解跨模態指令並輸出高品質立體聲影片成為業界關注的焦點。

🤔 背景或問題
先前的影片生成堆疊常被拆解為文字‑to‑影像、圖像‑to‑影像、首幀/末幀、主體參考、運動參考與影片編輯等多個專家模型,每個任務需要獨特的管線與模型。這種分散架構導致開發與維護成本升高,且難以用自然語言描述複雜的參考與編輯關係。

🧩 方法或架構
MiniMax H3 被定義為「通用多模態生成模型」,能同時讀取文字、圖像、影片與音訊作為單一統一的語境,並直接產出具備原聲立體聲的影片。核心技術包含:

  • Contextual Omni Representation:重新設計字幕生成,使其描述「來源內容與目標影片」之間的關係,而非只描述目標本身。
  • H3‑VAE:對詞彙表進行全面改造,高壓縮比帶來約 4× 的有效序列長度提升,降低訓練與推論成本,同時支援原生 2K 解析度。
  • H3‑Omni Transformer:顯著捨棄先前的 Hailuo‑02 架構,針對多模態語境導致的序列長度變異,將理解與生成工作負載分離,並依硬體特性調整各自的運算資源,據報告使端到端訓練吞吐量提升近 30%。
  • In‑Context Regeneration:不依賴外掛超解析度模型,基模型在原始多模態語境中自行重新生成低解析度輸出,從而在不猜測的情況下恢復細小文字與精細細節,對品牌與產品渲染尤為重要。

模型採用單一 API 端點,採用非同步三步驟流程:建立任務 → 輪詢 task_id → 下載 content.url。輸入端約需 100K tokens 的推論,經過內部蒸餾後平均約 4K tokens。

📊 數據或結果

  • 輸出規格:2K 解析度,影片長度 4–15 秒,僅支援整數秒數。
  • 訓練效能:端到端訓練吞吐量提升近 30%。
  • 成本聲稱:在 2K 下,每秒價格低於主流模型的三分之一;在 768p 下,低於主流 720p 的一半。
  • 第三方追蹤:2K 按使用計費約 0.13 美元/秒,約 1.95 美元/15 秒片段(僅供參考,未見於官方定價頁)。
  • 市場定位:根據 SCMP 引用的人工智慧分析,H3 在影片編輯領域領先,但在文字‑to‑影像方面落後於 Google Gemini Omni Flash,在 Flash,而在圖像‑to‑影像方面則落後於 Seedance 2.0 與 Gemini Omni Flash。

💡 深入分析
作者指出,語言成為橋梁:透過自然語言描述參考與編輯關係,將先前固定的任務集轉變為開放式、可描述的生成流程。這意味著開發者不再需要為每種子任務維護專門模型,而是透過同一端點調用不同的語境指令即可達成多樣化的影片產製需求。

⚠️ 限制

  • 目前僅透過平臺 API 與消費者 Hailuo AI App 使用,未提供自行部署的硬體方案。
  • 影片長度必須為整數秒(4、5、…、15 秒)。
  • 定價資訊主要來自第三方追蹤,官方頁面僅展示 Hailuo 2.3 級別,因此實際成本仍需以官方公告為準。

🎯 實務啟示
對於廣告、品牌、電商、產品設計、UI/UX、遊戲以及影片前視覺化等產業,MiniMax H3 提供一種「一端點多模態」的解決方案:透過簡單的文字敘述(例如「參考 Video 1 的鏡頭移動,讓 Image 2 中的角色唱歌,並將人聲與 Audio 3 對齊」),即可產出具備立體聲的 2K 影片,縮短從概念到成品的迭代週期。開發者可先註冊取得 API 金鑰,依照「建立任務 → 輪詢 → 下載」的流程整合至現有工作流程,並在成本模型上參考官方後續更新以評估是否符合預算。

🔗 來源

原始資料 MarkTechPost · 收集於 2026-08-02
來源原標題
MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio
作者
Asif Razzaq
原始標籤
AI Shorts · Applications · Artificial Intelligence · Audio Language Model · Editors Pick · Language Model · Large Language Model · Machine Learning · New Releases · Staff · Tech News · Technology · Voice AI
原始連結
https://www.marktechpost.com/2026/08/01/minimax-releases-minimax-h3-an-omni-modal-video-model-that-generates-15-second-2k-clips-with-native-stereo-audio/

摘要原文

MiniMax releases MiniMax H3 , a general-purpose multimodal generation model. MiniMax H3 is not a text-to-video model with add-ons. MiniMax describes it as a general-purpose multimodal generation model that reads text, images, video, and audio as one unified context and returns video with native stereo sound. The mains specs include: 2K output, 4–15 seconds, integer durations only . Previous video stacks split into text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference, and video editing, each often a separate expert model. MiniMax H3 folds those into one pretraining paradigm where reference and editing relationships are expressed in natural language. MiniMax’s example prompt makes the point: reference the camera movement from Video 1, have the character in Image 2 sing, match the vocals to Audio 3. Today: yes, through the API and no, on your own hardware. MiniMax launched H3 on July 31, 2026 with the model live in the platform API under the model ID MiniMax-H3 and in the consumer Hailuo AI app. Industries : MiniMax positions MiniMax H3 for advertising, branding, e-commerce, product design, UI/UX, and gaming along with film pre-visualization and retail catalog media. Applications : Ad variant generation, product and listing videos, animated posters, film title sequences, website hero loops, character-consistent game cinematics, and video-to-video motion transfer. The video generation guide documents three entry modes: text-to-video, first/last-frame image-to-video, and reference generation. Behind one endpoint and an asynchronous three-step flow: create a task, poll task_id , download content.url . Input limits worth designing around: Contextual Omni Representation : MiniMax rebuilt captioning so it describes the relationship between context and target video, not just the target. Most source material requires roughly 100K tokens of inference, distilled to about 4K tokens on average. Language is the bridge that turns a fixed task set into an open, descriptive one. H3-VAE : A full tokenizer overhaul. Its high compression ratio delivers a stated 4× gain in effective sequence length , cutting training and inference cost and it is the enabling technology for native 2K. H3-Omni Transformer : MiniMax explicitly set aside the Hailuo-02 architecture here. Multimodal context tripled sequence-length variance, so the training architecture separates understanding and generation workloads and tunes hardware utilization for each. Reported result: end-to-end training throughput up nearly 30% . In-Context Regeneration : Instead of a bolt-on super-resolution module, the base model regenerates its own low-resolution output in-context, re-reading the original multimodal context. That is what recovers small text and fine detail that conventional upscalers guess at — directly relevant to brand and product rendering. MiniMax’s own claim: at 2K, H3’s per-second price is less than a third of mainstream models; at 768p, less than half the price of mainstream 720p. The company amplified both the launch and the pricing framing on X ( 1 , 2 ). Third-party trackers and launch coverage put the 2K pay-as-you-go rate at $0.13 per second , about $1.95 for a 15-second clip, but MiniMax’s pay-as-you-go page still listed only Hailuo 2.3 tiers at the time of writing, so treat that figure as reported, not primary. On placement: SCMP reports , citing Artificial Analysis, that H3 leads in video editing while trailing Google’s Gemini Omni Flash in text-to-video and sitting behind both Seedance 2.0 and Gemini Omni Flash in image-to-video. Check out the Technical details . Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter . Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio appeared first on MarkTechPost .

tencent/hy3:free 自動生成