VoiceStudio
A local-first desktop studio for voice cloning, voice design, dubbing, dictation, transcription, and audiobook production.
VoiceStudio is an open-source desktop application for creating and processing speech on your own hardware. The README describes workflows for voice cloning, voice design, video dubbing, dictation, transcription, stories, and audiobooks, with 16 TTS engines, 11 ASR engines, and a catalogue covering up to 646 TTS languages depending on the selected engine. The default workflow is local: voices, projects, settings, and outputs stay on the machine, and the project does not require an account, API key, subscription, or usage meter for local use. It supports macOS Apple Silicon, Windows, Linux, and Docker, with CPU and optional GPU paths. The repository also documents a local REST, SSE, WebSocket, OpenAI-compatible audio API and an MCP server for integrations. The project is in active beta and the README recommends the latest release for stable work. First launch creates a managed Python environment and downloads the default model. Engine availability, language coverage, hardware requirements, and model or tokenizer terms vary, so the project documentation should be checked before production use. Voice cloning should be used only with the speaker's explicit permission.