Feature extraction AI Models

Models listed

50

Reviewed and live in this category

Worth weighing up

What goes in and out, licence terms, and whether you can host it yourself

How we treat data

Published specs only — we don't estimate numbers a vendor hasn't stated

Feature extraction models are a source-backed task collection in this directory. The current editorial inventory contains 25 mapped model records. Task membership is a discovery signal, not a performance ranking, deployment recommendation, or proof that every checkpoint accepts the same inputs and returns the same outputs.

Start with the model list, then open each record that matches your data, operating constraints, and intended workflow. Check the primary source, documented modalities, license, release information, and implementation notes before using a model in production.

This collection keeps unsupported pricing, availability, and benchmark claims blank. Each page remains outside the public index until an editor approves its content and related model records.

Feature extraction model list

Scan the table below, then open any model for pricing, limits and the full write-up.

50 models
ModelCapabilityContextInput $/1MOutput $/1MSpeed
Qwen/Qwen3-Embedding-8B
Qwen
22.232.8K$0.10$0.10
Qwen/Qwen3-Embedding-4B
Qwen
18.232K$0.00$0.00
Qwen/Qwen3-Embedding-0.6B
Qwen
11.08K$0.04$0.00
BAAI/bge-multilingual-gemma2
BAAI
10.58K$0.08$0.00
allenai/specter2_base
allenai
5.0
BAAI/bge-base-en
BAAI
5.0
BAAI/bge-base-en-v1.5
BAAI
5.0
BAAI/bge-base-zh-v1.5
BAAI
5.0
BAAI/bge-large-en-v1.5
BAAI
5.0
BAAI/bge-large-zh-v1.5
BAAI
5.0
BAAI/bge-reranker-large
BAAI
5.0
BAAI/bge-small-en
BAAI
5.0
BAAI/bge-small-en-v1.5
BAAI
5.0
BAAI/bge-small-zh-v1.5
BAAI
5.0
boboliu/Qwen3-Embedding-4B-W4A16-G128
boboliu
5.0
cambridgeltl/SapBERT-from-PubMedBERT-fulltext
cambridgeltl
5.0
facebook/encodec_24khz
facebook
5.0
facebook/hubert-base-ls960
facebook
5.0
facebook/w2v-bert-2.0
facebook
5.0
ibm-granite/granite-embedding-small-english-r2
ibm-granite
5.0
indobenchmark/indobert-base-p1
indobenchmark
5.0
intfloat/e5-mistral-7b-instruct
intfloat
5.0
intfloat/multilingual-e5-large
intfloat
5.0
intfloat/multilingual-e5-large-instruct
intfloat
5.0
jinaai/jina-clip-v2
jinaai
5.0
jinaai/jina-embeddings-v2-small-en
jinaai
5.0
jinaai/jina-embeddings-v3
jinaai
5.0
jinaai/jina-embeddings-v5-text-nano
jinaai
5.0
kyutai/mimi
kyutai
5.0
laion/clap-htsat-unfused
laion
5.0
laion/larger_clap_music_and_speech
laion
5.0
michaelfeil/bge-small-en-v1.5
michaelfeil
5.0
microsoft/wavlm-base-plus
microsoft
5.0
microsoft/wavlm-large
microsoft
5.0
mixedbread-ai/mxbai-embed-large-v1
mixedbread-ai
5.0
naver/splade-cocondenser-ensembledistil
naver
5.0
ncbi/MedCPT-Query-Encoder
ncbi
5.0
nvidia/llama-nemotron-embed-1b-v2
nvidia
5.0
OrdalieTech/Solon-embeddings-large-0.1
OrdalieTech
5.0
sonoisa/sentence-bert-base-ja-mean-tokens-v2
sonoisa
5.0
unslothai/1
unslothai
5.0
unslothai/lambda
unslothai
5.0
unslothai/other
unslothai
5.0
unslothai/repeat
unslothai
5.0
unslothai/vram-16
unslothai
5.0
WhereIsAI/UAE-Large-V1
WhereIsAI
5.0
Xenova/all-MiniLM-L6-v2
Xenova
5.0
Xenova/bge-base-en-v1.5
Xenova
5.0
Xenova/bge-small-en-v1.5
Xenova
5.0
YituTech/conv-bert-base
YituTech
5.0

50 models · click a column to sort

About the capability score: a 0–100 figure SyncDev calculates from each vendor's published specifications — context window, reasoning support, input modalities, tool calling, maximum output and how recently the model shipped. It measures breadth of capability, not benchmark performance, so a higher-scoring model is not automatically the better choice for your task.

Feature extraction model questions

It groups source-backed directory records for research and comparison. Review each model page and its primary documentation because collection membership alone does not prove quality, availability, or suitability.