About This Model
MISOTTS · Workflows
Miso Text-to-Speech S 8B is a text-to-speech model based on the Sesame CSM architecture. It generates Mimi audio codes from text and optional audio context, using a large Llama 3.2-style backbone and a smaller autoregressive audio decoder. The model is designed for high-quality conversational speech generation and voice continuation from prompt audio.
ModelTrends.ai Model ID
#28002
Flag content
Similar Models
No similar models found
AI Models for Image, Video & Text Generation
audio miso is a workflow for the MisoTTS family listed on ModelTrends.ai, a read-only catalog of open source AI image, video and text models. Compare it with other MisoTTS models, check its heat score to see how it is trending, and open it on PirateDiffusion or BitVector to try it.
Browse by Model Family
Text / LLM Models
411
Anima Models
509
Chroma Models
14
Flux Models
1,071
Flux 2 / Klein Models
1,237
MiniMax H3 Models
331
Hunyuan Models
197
Ideogram Models
20
Krea2 Models
674
Qwen Models
17
Qwen2 Models
36
Ltx2 Models
150
Stable Diffusion 1.5 Models
6,465
Stable Diffusion XL 1.0 Models
15,432
Zimage Models
804
Wan Models
919

Model
Trends
.ai
© ModelTrends.ai
|
Made in Japan
|
© 2026
