Audio-to-audio AI Models

Models listed

51

Reviewed and live in this category

Worth weighing up

What goes in and out, licence terms, and whether you can host it yourself

How we treat data

Published specs only — we don't estimate numbers a vendor hasn't stated

Audio-to-audio models are a source-backed task collection in this directory. The current editorial inventory contains 26 mapped model records. Task membership is a discovery signal, not a performance ranking, deployment recommendation, or proof that every checkpoint accepts the same inputs and returns the same outputs.

Start with the model list, then open each record that matches your data, operating constraints, and intended workflow. Check the primary source, documented modalities, license, release information, and implementation notes before using a model in production.

This collection keeps unsupported pricing, availability, and benchmark claims blank. Each page remains outside the public index until an editor approves its content and related model records.

Audio-to-audio model list

Scan the table below, then open any model for pricing, limits and the full write-up.

51 models
ModelCapabilityContextInput $/1MOutput $/1MSpeed
Gemini 3.5 Live Translate Preview
Google
39.1131.1K$3.50$21.00
Aratako/MioCodec-25Hz-24kHz
Aratako
10.0
Aratako/MioCodec-25Hz-44.1kHz-v2
Aratako
10.0
aufklarer/DeepFilterNet3-CoreML
aufklarer
10.0
aufklarer/PersonaPlex-7B-MLX-4bit
aufklarer
10.0
aufklarer/Sidon-CoreML
aufklarer
10.0
chenmozhijin/BSRoformer-GGUF
chenmozhijin
10.0
cstr/qwen3-tts-tokenizer-12hz-GGUF
cstr
10.0
HKUSTAudio/xcodec2
HKUSTAudio
10.0
JacobLinCool/MP-SENet-DNS
JacobLinCool
10.0
jhcodec/jhcodec
jhcodec
10.0
JorisCos/ConvTasNet_Libri1Mix_enhsingle_16k
JorisCos
10.0
JorisCos/ConvTasNet_Libri2Mix_sepclean_16k
JorisCos
10.0
JorisCos/DCCRNet_Libri1Mix_enhsingle_16k
JorisCos
10.0
julien-c/DPRNNTasNet-ks16_WHAM_sepclean
julien-c
10.0
JusperLee/Dolphin
JusperLee
10.0
JusperLee/TIGER-DnR
JusperLee
10.0
JusperLee/TIGER-speech
JusperLee
10.0
KRAFTON/Raon-SpeechChat-9B
KRAFTON
10.0
KrauthammerLab/cast-0.7b-s2s
KrauthammerLab
10.0
kyutai/hibiki-zero-3b-pytorch-bf16
kyutai
10.0
kyutai/moshika-rag-pytorch-bf16
kyutai
10.0
kyutai/personaplex-rl-seamless
kyutai
10.0
LiquidAI/LFM2.5-Audio-1.5B
LiquidAI
10.0
LiquidAI/LFM2.5-Audio-1.5B-JP-GGUF
LiquidAI
10.0
LocalAI-io/LocalVQE
LocalAI-io
10.0
lucadellalib/focalcodec_50hz
lucadellalib
10.0
lucadellalib/focalcodec_50hz_4k_causal
lucadellalib
10.0
microsoft/speecht5_vc
microsoft
10.0
mispeech/dashengtokenizer
mispeech
10.0
mpariente/DPRNNTasNet-ks2_WHAM_sepclean
mpariente
10.0
neuphonic/distill-neucodec
neuphonic
10.0
neuphonic/neucodec
neuphonic
10.0
neuphonic/neucodec-onnx-decoder
neuphonic
10.0
neuphonic/neucodec-onnx-decoder-int8
neuphonic
10.0
nvidia/bigvgan_22khz_80band
nvidia
10.0
nvidia/bigvgan_base_24khz_100band
nvidia
10.0
nvidia/bigvgan_v2_22khz_80band_256x
nvidia
10.0
nvidia/bigvgan_v2_22khz_80band_fmax8k_256x
nvidia
10.0
nvidia/bigvgan_v2_24khz_100band_256x
nvidia
10.0
nvidia/bigvgan_v2_44khz_128band_256x
nvidia
10.0
nvidia/bigvgan_v2_44khz_128band_512x
nvidia
10.0
nvidia/personaplex-7b-v1
nvidia
10.0
nvidia/RE-USE
nvidia
10.0
Qwen/Qwen3-TTS-Tokenizer-12Hz
Qwen
10.0
scragnog/Ace-Step-1.5-ScragVAE
scragnog
10.0
speechbrain/metricgan-plus-voicebank
speechbrain
10.0
speechbrain/sepformer-wham16k-enhancement
speechbrain
10.0
speechbrain/sepformer-whamr16k
speechbrain
10.0
speechbrain/sepformer-wsj02mix
speechbrain
10.0
tencent/Covo-Audio-Chat
tencent
10.0

51 models · click a column to sort

About the capability score: a 0–100 figure SyncDev calculates from each vendor's published specifications — context window, reasoning support, input modalities, tool calling, maximum output and how recently the model shipped. It measures breadth of capability, not benchmark performance, so a higher-scoring model is not automatically the better choice for your task.

Audio-to-audio model questions

It groups source-backed directory records for research and comparison. Review each model page and its primary documentation because collection membership alone does not prove quality, availability, or suitability.