Audio-text-to-text AI Models

Models listed

93

Reviewed and live in this category

Worth weighing up

What goes in and out, licence terms, and whether you can host it yourself

How we treat data

Published specs only — we don't estimate numbers a vendor hasn't stated

Audio-text-to-text models are a source-backed task collection in this directory. The current editorial inventory contains 68 mapped model records. Task membership is a discovery signal, not a performance ranking, deployment recommendation, or proof that every checkpoint accepts the same inputs and returns the same outputs.

Start with the model list, then open each record that matches your data, operating constraints, and intended workflow. Check the primary source, documented modalities, license, release information, and implementation notes before using a model in production.

This collection keeps unsupported pricing, availability, and benchmark claims blank. Each page remains outside the public index until an editor approves its content and related model records.

Audio-text-to-text model list

Scan the table below, then open any model for pricing, limits and the full write-up.

93 models
ModelCapabilityContextInput $/1MOutput $/1MSpeed
Gemini 3.5 Flash Lite
Google
89.41.0M$0.30$2.50
Gemini 3.6 Flash
Google
89.41.0M$1.50$7.50
Gemini 3.1 Flash Lite
Google
89.01.0M$0.25$1.50
Gemini 3.5 Flash
Google
89.01.0M$1.50$9.00
Gemini Flash Latest
Google
89.01.0M$1.50$9.00
Gemini Flash-Lite Latest
Google
89.01.0M$0.25$1.50
Deep Research Max Preview
Google
88.91.0M$2.00$12.00
Gemini Deep Research Preview
Google
88.91.0M$2.00$12.00
Gemini 3.1 Flash Lite Preview
Google
88.61.0M$0.25$1.50
Gemini 3.1 Pro Preview
Google
88.51.0M$2.00$12.00
Gemini 3.1 Pro Preview Custom Tools
Google
88.51.0M$2.00$12.00
Gemini 3 Flash Preview
Google
88.21.0M$0.50$3.00
Gemini 3 Pro Preview
Google
88.01.0M$2.00$12.00
Gemini 2.5 Flash
Google
87.21.0M$0.30$2.50
Gemini 2.5 Flash-Lite
Google
87.21.0M$0.10$0.40
Gemini 2.5 Pro
Google
87.21.0M$1.25$10.00
MiMo-V2.5
Xiaomi
80.41.0M$0.14$0.28
MiMo-V2-Omni
Xiaomi
78.5262.1K$0.14$0.28
Qwen3.6 27B
Alibaba
77.2262.1K$0.60$3.60
Qwen3.6 35B-A3B
Alibaba
77.2262.1K$0.25$1.49
Qwen3.5 122B-A10B
Alibaba
76.9262.1K$0.40$3.20
Qwen3.5 27B
Alibaba
76.9262.1K$0.30$2.40
Qwen3.5 35B-A3B
Alibaba
76.9262.1K$0.25$2.00
Inkling
Thinkingmachines
76.81.0M$1.00$4.05
Inkling Small
Thinkingmachines
76.81.0M$0.50$1.20
Qwen3.5 397B-A17B
Alibaba
76.8262.1K$0.60$3.60
Gemini Robotics-ER 1.6 Preview
Google
73.8131.1K$1.00$5.00
Gemini 3.1 Flash Live Preview
Google
73.7131.1K$0.75$4.50
Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA
72.1256K$0.00$0.00
GPT-Realtime-2.1
OpenAI
67.6128K$4.00$24.00
Gemini 2.0 Flash
Google
66.61.0M$0.10$0.40
Gemini 2.0 Flash-Lite
Google
66.61.0M$0.07$0.30
Gemma 4 E2B IT
Google
64.2131.1K$0.02$0.10
Gemma 4 E4B IT
Google
64.2131.1K$0.04$0.20
Gemini 1.5 Pro
Google
63.62M$1.25$5.0065 t/s
GPT-4o
OpenAI
46.8128K$2.50$10.00110 t/s
Qwen-Omni Turbo
Alibaba
42.132.8K$0.07$0.27
Gemini 3.5 Live Translate Preview
Google
39.1131.1K$3.50$21.00
Nemotron VoiceChat
NVIDIA
39.0128K$0.00$0.00
mistralai/Voxtral-Small-24B-2507
mistralai
32.632.8K$0.23$0.51
Gemini Embedding 2
Google
28.88.2K$0.20$0.00
GPT Realtime Whisper
OpenAI
9.50
Whisper 3 Large
OpenAI
7.9448$0.00$0.00
Whisper Large v3 Turbo
OpenAI
6.3448
ACE-Step/acestep-transcriber
ACE-Step
5.0
ai-babai/gigachat-audio-mlx-q8-bf16
ai-babai
5.0
bartowski/mistralai_Voxtral-Mini-3B-2507-GGUF
bartowski
5.0
bartowski/mistralai_Voxtral-Small-24B-2507-GGUF
bartowski
5.0
cstr/MOSS-Audio-4B-Instruct-GGUF
cstr
5.0
DeSTA-ntu/DeSTA2.5-Audio-Llama-3.1-8B
DeSTA-ntu
5.0
fixie-ai/ultravox-v0_2
fixie-ai
5.0
fixie-ai/ultravox-v0_3
fixie-ai
5.0
fixie-ai/ultravox-v0_4
fixie-ai
5.0
fixie-ai/ultravox-v0_5-llama-3_1-8b
fixie-ai
5.0
fixie-ai/ultravox-v0_5-llama-3_2-1b
fixie-ai
5.0
fixie-ai/ultravox-v0_5-llama-3_3-70b
fixie-ai
5.0
fixie-ai/ultravox-v0_6-gemma-3-27b
fixie-ai
5.0
fixie-ai/ultravox-v0_6-llama-3_1-8b
fixie-ai
5.0
fixie-ai/ultravox-v0_6-llama-3_3-70b
fixie-ai
5.0
fixie-ai/ultravox-v0_6-qwen-3-32b
fixie-ai
5.0
fixie-ai/ultravox-v0_7-glm-4_6
fixie-ai
5.0
ggml-org/ultravox-v0_5-llama-3_2-1b-GGUF
ggml-org
5.0
IHP-Lab/Qwen2-Audio_PCLM_DPO
IHP-Lab
5.0
microsoft/VibeVoice-ASR-HF
microsoft
5.0
mispeech/GLAP
mispeech
5.0
mispeech/midashenglm-7b-0804-fp32
mispeech
5.0
mlx-community/MOSS-Music-8B-Thinking-8bit
mlx-community
5.0
mlx-community/Qwen2-Audio-7B-Instruct-4bit
mlx-community
5.0
MohamedRashad/Voxtral-Mini-3B-2507-transformers
MohamedRashad
5.0
MohamedRashad/Voxtral-Small-24B-2507-transformers
MohamedRashad
5.0
mradermacher/Ichigo-llama3.1-s-instruct-v0.4-i1-GGUF
mradermacher
5.0
NexaAI/OmniAudio-2.6B
NexaAI
5.0
NexaAI/Qwen2-Audio-7B-GGUF
NexaAI
5.0
nvidia/audio-flamingo-3-hf
nvidia
5.0
nvidia/audio-flamingo-next-captioner-hf
nvidia
5.0
nvidia/audio-flamingo-next-hf
nvidia
5.0
nvidia/audio-flamingo-next-think-hf
nvidia
5.0
nvidia/music-flamingo-2601-hf
nvidia
5.0
nvidia/music-flamingo-hf
nvidia
5.0
nvidia/music-flamingo-think-2601-hf
nvidia
5.0
onnx-community/Voxtral-Mini-3B-2507-ONNX
onnx-community
5.0
OpenMOSS-Team/MOSS-Audio-4B-Instruct
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Audio-4B-Thinking
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Audio-8B-Instruct
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Audio-8B-Thinking
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Music-8B-Instruct
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Music-8B-Thinking
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-Transcribe-Diarize
OpenMOSS-Team
5.0
Qwen/Qwen2-Audio-7B
Qwen
5.0
Qwen/Qwen2-Audio-7B-Instruct
Qwen
5.0
sarvamai/shuka-1
sarvamai
5.0
second-state/Qwen2-Audio-7B-Instruct-GGUF
second-state
5.0
software-mansion/react-native-executorch-gemma-4-multimodal
software-mansion
5.0

93 models · click a column to sort

About the capability score: a 0–100 figure SyncDev calculates from each vendor's published specifications — context window, reasoning support, input modalities, tool calling, maximum output and how recently the model shipped. It measures breadth of capability, not benchmark performance, so a higher-scoring model is not automatically the better choice for your task.

Audio-text-to-text model questions

It groups source-backed directory records for research and comparison. Review each model page and its primary documentation because collection membership alone does not prove quality, availability, or suitability.