Voice Interfaces for Everyone
Voice is an open source AI toolkit for developers building real-time voice agents and applications.
Full guides, models, and API reference are at
moonshine-voice.readthedocs.io
.
Everything runs on-device — fast, private, and with no account or API keys.
Optimized for live streaming, with low latency by doing work while the user is still talking.
Speech to text models trained from scratch, from
higher accuracy than Whisper Large V3
down to
.
One library across
Python, JavaScript/WASM, iOS, Android, macOS, Linux, Windows, and Raspberry Pi
.
Quickstart
pip install moonshine-voice moonshine-voice mic --language enEvery other platform is covered in the
, with runnable samples in
.
Building with a coding agent? Copy
.agents/skills/moonshine-voice/
into your project's .agents/skills/ folder, or run npx skills add moonshine-ai/moonshine --skill moonshine-voice. The skill teaches the current API shape so the agent does not reach for Whisper or the old DialogFlow names.
Documentation
to install and run on your platform.
for transcription, text to speech, and conversational agents.
for what's available, accuracy, and domain customization.
for the classes, options, and C API.
Support
for live support.
for bugs and feature requests.
License
Licensed under the
. The models are MIT by default too, in every language and at every size — the only exceptions are the legacy non-streaming models for languages other than English, which stay under the non-commercial Moonshine Community License and are enumerated in
.