Audio · Speech-to-text · Released Sep 2026
Meta Muse Voice Transcribe 1.0
Meta's speech-to-text model for push-to-talk and speaker-aware transcription, which returns plain text.
From Meta, Muse Voice Transcribe 1.0 is a cloud speech-to-text model. It transcribes spoken audio into written text. It runs through OpenRouter using your own API key, from about $0.003 per minute of audio.
Specs
- Released
- this month (Sep 2026)
- Pricing
- $0.003 per minute of audio
- Type
- Speech-to-text
About the creator
Meta
Meta's FAIR and GenAI labs build the open-weight Llama family of language models — among the most widely used open models for local inference.
ai.meta.com ↗More from Meta
Muse Image
ImageCloud
Meta Superintelligence Labs' agentic image model, which plans a layout and calls search and coding tools before it renders, for prompt-faithful generation, multi-reference composition and precise local edits.
Llama 3.2 1B
TextLocal
Meta's 1B instruction-tuned model — the smallest Llama, runnable in-process via HuggingFace or through Ollama.
Llama 3.1 8B
TextLocal
Meta's 8B instruction-tuned model with 128K context and strong general reasoning.
Llama 3.2 3B
TextLocal
Meta's compact 3B instruction-tuned model — fast on-device text generation.
Muse Glimmer
TextLocal
Meta Superintelligence Labs' 30B multimodal model for always-on local agents — tuned for tool use, long tasks, and failure recovery.