AudioLocal AI

Miso One (Miso TTS)

Open-weights model (released under a modified MIT license), free to download and run yourself; you provide your own compute. Miso Labs has indicated hosted API access is coming. Check the official sources for details. Check official pricing →

At a Glance

Miso One (Miso TTS) is an open-weights text-to-speech model from Miso Labs, released in 2026. It’s an 8-billion-parameter model focused on expressive, emotive speech, and its weights are published on Hugging Face under a modified MIT license. It’s aimed more at developers and voice-AI tinkerers than at non-technical users. It shines if you want to run an open voice model yourself rather than rent a hosted service.

What Miso One Is Best For

  • Voice generation: expressive, conversational AI speech
  • AI narration: voiceovers for videos, demos, and prototypes
  • Developers: building voice features on an open-weights model
  • Local and open-model workflows: running TTS on your own hardware
  • Audio prototypes and accessibility: experimenting with voice interfaces

How Miso One Works

Miso TTS generates speech from text and can also condition on a short audio sample to match a speaker’s tone. It’s built on an open architecture and distributed as downloadable weights, so you run it in your own environment (locally or on your own server). Audio is watermarked by default.

Getting Better Results

Provide clean reference audio. When matching a voice, a clear, short sample tends to produce better tone matching.

Mind your hardware. An 8B model needs a capable machine; check the model card for requirements before committing.

Iterate on phrasing. Punctuation and phrasing affect expressiveness, so small text edits can noticeably change delivery.

Honest Limitations

  • More technical than consumer voice apps: setup assumes some comfort with models and hardware
  • Ethical and permission concerns: voice cloning can be misused; only clone voices you’re authorized to use
  • Evolving availability: hosted API and tooling are still developing

Safety note: Only clone or generate voices you have explicit permission to use. Misusing someone’s voice can be harmful and may be illegal.

Alternatives Worth Knowing

  • ElevenLabs, leading hosted, easy-to-use voice generation
  • Descript, audio/video editing with AI voice features
  • Hugging Face, where you can find and run open models like this one
  • Suno, AI music generation (a related but distinct audio use case)
  • Udio, AI music creation

Continue learning

Explore related guides, tools, workflows, and prompts that help you go deeper into this topic.

See how this tool fits into a workflow

Browse step-by-step AI workflows that use ChatGPT, Claude, Gemini, and other tools.

Frequently Asked Questions

What is Miso One / Miso TTS?

Miso One is the name people use for Miso Labs' Miso TTS model, an 8-billion-parameter, open-weights text-to-speech system focused on expressive, emotive English speech. The weights are published on Hugging Face under a modified MIT license, so developers can download and run it themselves.

Is Miso TTS free to use?

The model weights are open and free to download and run on your own hardware. You cover your own compute, and Miso Labs has signaled that a hosted API is coming. Always check the current license terms before commercial use.

Can Miso TTS clone voices?

Miso TTS can condition on audio context and reportedly supports one-shot voice cloning from a short clip. Only clone or generate voices you have clear permission to use. Generated audio is watermarked by default, but you remain responsible for using voice cloning ethically and legally.

Last updated: