Moshi
Real-time voice AI by Kyutai that thinks while speaking with 160ms latency - free open-source speech model trained without scraped human voice data.
Moshi is a real-time speech-to-speech conversational AI developed by Kyutai, a French non-profit AI research lab, and released publicly in June 2024 as both a live demo at moshi.chat and an open-source model under a CC BY 4.0 license. Unlike GPT-4o voice or ElevenLabs, Moshi operates as a single end-to-end audio model rather than a speech-to-text, LLM, and text-to-speech pipeline, enabling it to generate audio tokens in parallel with listening and processing and achieving 160ms theoretical latency compared to 300-500ms typical of pipeline architectures. Moshi uses a dual-stream architecture that models both the AI inner monologue and the spoken output simultaneously, allowing it to continue listening and adjust responses while speaking rather than waiting for a full turn to complete. The model weights and training code are released under CC BY 4.0, making Moshi the first fully open-source real-time conversational speech model released by a major research lab.
Key Features
- End-to-end audio token model processes audio in and out without intermediate speech-to-text conversion, enabling 160ms theoretical latency versus 300-500ms for pipeline approaches
- Dual-stream architecture models the AI inner monologue and spoken output simultaneously, allowing the model to think and speak in parallel rather than sequentially
- Real-time interruption handling detects and responds to user interjections mid-sentence without waiting for the AI to finish its current utterance
- Open-source model weights released under CC BY 4.0 allow researchers and developers to fine-tune, self-host, and build on Moshi without commercial licensing restrictions
- Live web demo at moshi.chat lets anyone try real-time voice conversation directly in a browser without signup, API keys, or software installation
- Trained without scraped human voice data - Moshi uses a synthetic voice based on a consented voice actor rather than internet audio, avoiding speaker consent issues
- Multi-stream audio architecture supports overlapping speech, background noise awareness, and turn-taking cues that improve over single-channel voice models
Use Cases
- AI researchers studying real-time speech interaction architectures who need an open-source baseline to compare against proprietary voice models from OpenAI and Google
- Developers building voice-first applications who want to self-host a low-latency speech model without licensing fees or dependency on commercial voice APIs
- Linguists and interaction researchers using the open weights to fine-tune conversational turn-taking behavior on specific languages or interaction styles
- Hobbyists and developers exploring voice AI building local demos and prototypes with Moshi weights running on personal hardware without API costs
Pros
- 160ms theoretical latency from the end-to-end audio architecture is the lowest publicly documented latency of any conversational voice AI model as of mid-2024
- CC BY 4.0 open weights with public training code make Moshi the only fully open conversational voice model - researchers can study and improve the architecture directly
- No scraped human voice data in training addresses voice consent issues that affect models trained on internet audio recordings without speaker permission
Cons
- Current conversation quality and knowledge accuracy is below GPT-4o voice and Gemini Live - Moshi is a research model, not production-ready for customer-facing applications
- Self-hosting requires at least 24GB VRAM for real-time inference - consumer hardware cannot run Moshi at the documented 160ms latency target
- English-only at public release - the multilingual conversational speech market requires separate fine-tuned checkpoints not yet publicly available
Moshi Alternatives
Explore similar tools and alternatives
Looking for alternatives to Moshi? Here are some similar tools you might like:
ElevenLabs
AI voice generation platform for creating realistic text-to-speech, voice cloning, and multilingual dubbing.
Hume AI
Empathic AI voice API that measures and responds to human emotion in real time - powering emotionally intelligent voice assistants via the EVI 2 interface.
Vapi
Voice AI infrastructure API for building AI phone agents - handles call routing, real-time transcription, voice synthesis, and LLM orchestration for developers.
Ready to try Moshi?
Visit the official website to explore all features and get started with Moshi today.
Reviews
0 reviews for Moshi
Based on 0 reviews
Share your experience
Log in to write a review for Moshi
ElevenLabs
AI voice generation platform for creating realistic text-to-speech, voice cloning, and multilingual dubbing.
Murf AI
AI voice generator with 200+ voices in 35+ languages, 55ms API latency, and AI dubbing with lip sync.
Suno AI
AI music generator that creates full songs with vocals, instruments, and lyrics from a text prompt in seconds.
Have an AI Tool?
List your AI tool for free, or go featured for top placement in your category - and reach thousands of potential users.
Submit Your Tool