Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human
https://techcrunch.com/2026/07/31/smallest-ai-raises-13m-to-build-ultra-fast-voice-ai-that-sounds-genuinely-human/📌 【Smallest.ai 融資新聞】不追求大模型規模,專注「邊聽、邊想、邊說」的超低延遲語音 AI
TL;DR:Smallest.ai 獲 1300 萬美元融資,致力於開發能實現「人機難以分辨」的超快速語音模型。
當前的 AI Agent 雖然能解決客戶服務問題,但對話中的停頓與機械感,仍讓使用者一眼看穿對方是機器。Smallest.ai 認為,語音 Agent 的下一個突破點不在於把大型語言模型(LLM)做得更快,而在於使用專為人類對話設計的「小型化、專業化」模型。
🤔 打破 LLM 的思考延遲:模擬人類的即時處理能力
目前的 LLM 運作邏輯通常是接收完整提示詞(Prompt)後才開始思考,這種延遲在文字對話中尚可接受,但在語音對話中,即便是短暫的停頓也會讓對話顯得不自然。
Smallest.ai 的設計理念是模仿人類的資訊處理方式:在說話的同時,大腦就已經在進行聽覺接收、思考與語言產出。這種設計讓模型能像真人一樣,在對話過程中處理資訊,而非等待整段音訊結束後才反應。
🧩 雙模型架構:小型語音模型 + 大型離線 LLM
Smallest.ai 預期未來的 AI Agent 將依賴兩種模型協同工作:
- 小型語音模型 (Small Voice Model):負責即時互動,處理語音特有的細微差別,如多種口音、數十種語言,以及在嘈熱環境下的運作。
- 大型基礎模型 (Large Foundational Model):當小型模型遇到知識範圍外的複雜問題時,會將請求移交給大型模型,並像真人一樣讓客戶「稍等片刻」以進行研究。
📊 專注於實時對話,而非影音剪輯
與 ElevenLabs、Cartesia 等將語音 AI 應用於音訊配音或 Podcast 的競品不同,Smallest.ai 嚴格專注於企業級客戶的「實時對話型語音 Agent」。
其目前的客戶已包含 RingCentral 與 Truecaller。該公司認為,對於專注於客戶服務的初創公司來說,開發高品質的語音模型會分散其核心業務的注意力,而 Smallest.ai 正是為了填補這項技術缺口。
🎯 目標是突破圖靈測試
Smallest.ai 的終極目標是讓對話品質達到「無法分辨是 AI 還是真人」的程度。透過將即時互動的延遲降至趨近於零,他們希望打破語音通訊中的機器感。
🔗 來源
- 標題:Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human
- 作者/機構:Marina Temkin @ TechCrunch
- 連結:https://techcrunch.com/2026/07/31/smallest-ai-raises-13m-to-build-ultra-fast-voice-ai-that-sounds-genuinely-human/
#AI #VoiceAI #SmallestAI #LLM #MachineLearning #CustomerService #TechNews #ArtificialIntelligence #SpeechTechnology #Startup
原始資料 TechCrunch AI · 收集於 2026-08-01
摘要原文
While AI agents are increasingly capable of solving customer support problems, most people can still tell immediately when they’re talking to a machine instead of a human. Smallest.ai , a startup founded in late 2024, is betting the next leap in voice agents will not come from making large language models faster, but from using smaller, specialized models built for human conversation. Simply put, the company wants to make speaking to an AI agent indistinguishable from talking to a human. To do so, it’s developing a small voice model designed to mimic how humans process information by listening, thinking, and speaking simultaneously. “While I’m speaking to you, you’re already thinking, and you might interrupt me if I talk for too long,” Sudarshan Kamath (pictured left), founder and CEO of Smallest.ai, told TechCrunch, adding that this is exactly how the startup’s model is designed to work. To fuel this mission, Smallest.ai has raised $13 million in a Series A round, led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital. The fresh capital brings the startup’s total funding to over $21 million. “The way an LLM works is you give it an entire prompt, and then it starts thinking,” Kamath said. While that latency is acceptable in a text chat, in a voice conversation, even a short pause feels unnatural. “If you think about how we are talking, I’m not giving you like a large clipping of my audio, and then you start thinking.” The startup’s model serves as a real-time intelligence layer that enables natural customer conversations on specific topics, with virtually zero response lag. But if the model encounters a subject outside its limited knowledge base, Smallest.ai hands off the query to a large foundational model, briefly placing the customer on hold to “research” the issue — just as a real human would do. Kamath believes that all AI agents will soon rely on two models: a small voice model for real-time interaction, and an “offline” LLM that is called upon as needed to solve complex problems. Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances, such as handling diverse accents, supporting dozens of languages, and operating in noisy environments. The startup’s existing customers include companies in the voice space, including RingCentral and Truecaller. Kamath said that any customer support company, including newer ones like Sierra and Decagon, is a potential customer for the startup. When asked why a well-funded AI customer support company wouldn’t build its own voice model, Kamath said that for customer support startups, becoming “extremely good at doing voice is a distraction from their core business.” Smallest.ai competes with voice AI leader ElevenLabs, as well as Cartesia and regional players like Sarvam that focus on local languages. While some competitors apply voice AI to use cases, like audio dubbing and podcasting, Smallest.ai focuses strictly on real-time conversational voice agents for its enterprise customers. “We want our models to break the Turing test,” Kamath said. “You should speak to our model and not know it’s AI or human. That’s the sole focus of the company.”
由 tencent/hy3:free 自動生成