Audio classification AI Models

Models listed

50

Reviewed and live in this category

Worth weighing up

What goes in and out, licence terms, and whether you can host it yourself

How we treat data

Published specs only — we don't estimate numbers a vendor hasn't stated

Audio classification models are a source-backed task collection in this directory. The current editorial inventory contains 25 mapped model records. Task membership is a discovery signal, not a performance ranking, deployment recommendation, or proof that every checkpoint accepts the same inputs and returns the same outputs.

Start with the model list, then open each record that matches your data, operating constraints, and intended workflow. Check the primary source, documented modalities, license, release information, and implementation notes before using a model in production.

This collection keeps unsupported pricing, availability, and benchmark claims blank. Each page remains outside the public index until an editor approves its content and related model records.

Audio classification model list

Scan the table below, then open any model for pricing, limits and the full write-up.

50 models
ModelCapabilityContextInput $/1MOutput $/1MSpeed
alefiury/wav2vec2-large-xlsr-53-gender-recognition-librispeech
alefiury
10.0
anton-l/wav2vec2-random-tiny-classifier
anton-l
10.0
audeering/wav2vec2-large-robust-12-ft-emotion-msp-dim
audeering
10.0
audeering/wav2vec2-large-robust-24-ft-age-gender
audeering
10.0
audeering/wav2vec2-large-robust-6-ft-age-gender
audeering
10.0
aufklarer/Qwen3-ForcedAligner-0.6B-4bit
aufklarer
10.0
aufklarer/Silero-VAD-v6.2.1-MLX
aufklarer
10.0
aufklarer/WeSpeaker-ResNet34-LM-MLX
aufklarer
10.0
awsaf49/sonics-spectttra-alpha-120s
awsaf49
10.0
awsaf49/sonics-spectttra-beta-5s
awsaf49
10.0
awsaf49/sonics-spectttra-gamma-5s
awsaf49
10.0
Dpngtm/wav2vec2-emotion-recognition
Dpngtm
10.0
ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition
ehcalabres
10.0
facebook/audiobox-aesthetics
facebook
10.0
facebook/mms-lid-1024
facebook
10.0
facebook/mms-lid-126
facebook
10.0
facebook/mms-lid-256
facebook
10.0
facebook/mms-lid-4017
facebook
10.0
firdhokk/speech-emotion-recognition-with-openai-whisper-large-v3
firdhokk
10.0
griko/gender_cls_svm_ecapa_voxceleb
griko
10.0
JaesungHuh/voice-gender-classifier
JaesungHuh
10.0
jakeBland/wav2vec-vm-finetune
jakeBland
10.0
justin1983/wav2vec2-base-finetuned-amd
justin1983
10.0
Jzuluaga/accent-id-commonaccent_ecapa
Jzuluaga
10.0
laion/clap-htsat-fused
laion
10.0
m-a-p/MERT-v1-330M
m-a-p
10.0
m-a-p/MERT-v1-95M
m-a-p
10.0
MelodyMachine/Deepfake-audio-detection-V2
MelodyMachine
10.0
mispeech/ced-base
mispeech
10.0
mispeech/ced-tiny
mispeech
10.0
MIT/ast-finetuned-audioset-10-10-0.4593
MIT
10.0
MIT/ast-finetuned-audioset-14-14-0.443
MIT
10.0
mo-thecreator/Deepfake-audio-detection
mo-thecreator
10.0
mudler/ced-gguf
mudler
10.0
onecxi/open-vakgyata
onecxi
10.0
OpenMuQ/MuQ-large-msd-iter
OpenMuQ
10.0
OpenMuQ/MuQ-MuLan-large
OpenMuQ
10.0
prithivMLmods/Common-Voice-Gender-Detection
prithivMLmods
10.0
prithivMLmods/Common-Voice-Gender-Detection-ONNX
prithivMLmods
10.0
Speech-Arena-2025/DF_Arena_500M_V_1
Speech-Arena-2025
10.0
speechbrain/emotion-recognition-wav2vec2-IEMOCAP
speechbrain
10.0
speechbrain/lang-id-voxlingua107-ecapa
speechbrain
10.0
speechbrain/spkrec-xvect-voxceleb
speechbrain
10.0
StanislavKo28/music_moods_classification
StanislavKo28
10.0
superb/hubert-base-superb-er
superb
10.0
superb/hubert-large-superb-er
superb
10.0
superb/wav2vec2-base-superb-er
superb
10.0
superb/wav2vec2-base-superb-ks
superb
10.0
thelou1s/yamnet
thelou1s
10.0
xbgoose/hubert-large-speech-emotion-recognition-russian-dusha-finetuned
xbgoose
10.0

50 models · click a column to sort

About the capability score: a 0–100 figure SyncDev calculates from each vendor's published specifications — context window, reasoning support, input modalities, tool calling, maximum output and how recently the model shipped. It measures breadth of capability, not benchmark performance, so a higher-scoring model is not automatically the better choice for your task.

Audio classification model questions

It groups source-backed directory records for research and comparison. Review each model page and its primary documentation because collection membership alone does not prove quality, availability, or suitability.