Video-text-to-text AI Models

Models listed

78

Reviewed and live in this category

Worth weighing up

What goes in and out, licence terms, and whether you can host it yourself

How we treat data

Published specs only — we don't estimate numbers a vendor hasn't stated

Video-text-to-text models are a source-backed task collection in this directory. The current editorial inventory contains 78 mapped model records. Task membership is a discovery signal, not a performance ranking, deployment recommendation, or proof that every checkpoint accepts the same inputs and returns the same outputs.

Start with the model list, then open each record that matches your data, operating constraints, and intended workflow. Check the primary source, documented modalities, license, release information, and implementation notes before using a model in production.

This collection keeps unsupported pricing, availability, and benchmark claims blank. Each page remains outside the public index until an editor approves its content and related model records.

Video-text-to-text model list

Scan the table below, then open any model for pricing, limits and the full write-up.

78 models
ModelCapabilityContextInput $/1MOutput $/1MSpeed
Gemini 3.5 Flash Lite
Google
89.41.0M$0.30$2.50
Gemini 3.6 Flash
Google
89.41.0M$1.50$7.50
Gemini 3.1 Flash Lite
Google
89.01.0M$0.25$1.50
Gemini 3.5 Flash
Google
89.01.0M$1.50$9.00
Gemini Flash Latest
Google
89.01.0M$1.50$9.00
Gemini Flash-Lite Latest
Google
89.01.0M$0.25$1.50
Deep Research Max Preview
Google
88.91.0M$2.00$12.00
Gemini Deep Research Preview
Google
88.91.0M$2.00$12.00
Gemini 3.1 Flash Lite Preview
Google
88.61.0M$0.25$1.50
Gemini 3.1 Pro Preview
Google
88.51.0M$2.00$12.00
Gemini 3.1 Pro Preview Custom Tools
Google
88.51.0M$2.00$12.00
Gemini 3 Flash Preview
Google
88.21.0M$0.50$3.00
Gemini 3 Pro Preview
Google
88.01.0M$2.00$12.00
Gemini 2.5 Flash
Google
87.21.0M$0.30$2.50
Gemini 2.5 Flash-Lite
Google
87.21.0M$0.10$0.40
Gemini 2.5 Pro
Google
87.21.0M$1.25$10.00
Muse Spark 1.1
Meta
82.01M$1.25$4.25
Kimi K3
Moonshot AI
80.81.0M$3.00$15.00
MiMo-V2.5
Xiaomi
80.41.0M$0.14$0.28
Qwen3.7 Flash
Alibaba
79.11M$0.03$0.13
Qwen3.6 Flash
Alibaba
78.71M$0.19$1.13
MiMo-V2-Omni
Xiaomi
78.5262.1K$0.14$0.28
Qwen3.6 27B
Alibaba
77.2262.1K$0.60$3.60
Qwen3.6 35B-A3B
Alibaba
77.2262.1K$0.25$1.49
Qwen3.5 122B-A10B
Alibaba
76.9262.1K$0.40$3.20
Qwen3.5 27B
Alibaba
76.9262.1K$0.30$2.40
Qwen3.5 35B-A3B
Alibaba
76.9262.1K$0.25$2.00
Qwen3.5 397B-A17B
Alibaba
76.8262.1K$0.60$3.60
Qwen3.8 Max
Alibaba
75.71M$2.00$6.00
Qwen3.8 Max Preview
Alibaba
75.61M$1.50$5.00
Kimi K2.7 Code
Moonshot AI
74.9262.1K$0.95$4.00
Kimi K2.7 Code Highspeed
Moonshot AI
74.9262.1K$1.90$8.00
Kimi K2.6
Moonshot AI
74.6262.1K$0.95$4.00
Kimi K2.5
Moonshot AI
74.0262.1K$0.60$3.00
Gemini Robotics-ER 1.6 Preview
Google
73.8131.1K$1.00$5.00
Qwen3.7 Plus
Alibaba
73.81M$0.50$3.00
Gemini 3.1 Flash Live Preview
Google
73.7131.1K$0.75$4.50
Qwen3.6 Plus
Alibaba
73.51M$0.50$3.00
Qwen3.5 Plus
Alibaba
73.31M$0.40$2.40
GLM-5V-Turbo
Zhipu AI
72.3200K$5.00$22.00
MiniMax-M3
MiniMax
72.1512K$0.30$1.20
Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA
72.1256K$0.00$0.00
Qwen3.5 9B
Alibaba
71.9262.1K$0.04$0.15
Step 3.7 Flash
StepFun
69.7256K$0.18$1.11
Gemini 2.0 Flash
Google
66.61.0M$0.10$0.40
Gemini 2.0 Flash-Lite
Google
66.61.0M$0.07$0.30
Nemotron Nano 12B v2 VL
NVIDIA
64.2128K$0.20$0.60
Gemini 1.5 Pro
Google
63.62M$1.25$5.0065 t/s
Nano Banana 2
Google
62.5131.1K$0.50$60.00
GLM-4.6V
Zhipu AI
61.5128K$0.30$0.90
GLM-4.5V
Zhipu AI
56.064K$0.60$1.80
Qwen-Omni Turbo
Alibaba
42.132.8K$0.07$0.27
Gemini Embedding 2
Google
28.88.2K$0.20$0.00
allenai/MolmoPoint-Vid-4B
allenai
5.0
chenjoya/videollm-online-8b-v1plus
chenjoya
5.0
DAMO-NLP-SG/VideoLLaMA3-2B
DAMO-NLP-SG
5.0
DAMO-NLP-SG/VideoLLaMA3-7B
DAMO-NLP-SG
5.0
Diankun/Spatial-MLLM-v1.1-Instruct-135K
Diankun
5.0
jdopensource/JoyAI-VL-Interaction
jdopensource
5.0
jdopensource/JoyAI-VL-Interaction-Preview
jdopensource
5.0
Kwai-Keye/Keye-VL-1_5-8B
Kwai-Keye
5.0
Kwai-Keye/Keye-VL-8B-Preview
Kwai-Keye
5.0
llava-hf/LLaVA-NeXT-Video-7B-hf
llava-hf
5.0
lmms-lab/LLaVA-Video-7B-Qwen2
lmms-lab
5.0
MCG-NJU/TimeLens2-8B
MCG-NJU
5.0
MCG-NJU/VideoChat3-4B
MCG-NJU
5.0
mlx-community/SmolVLM2-500M-Video-Instruct-mlx
mlx-community
5.0
NemoStation/Marlin-2B
NemoStation
5.0
nvidia/nemotron-labs-audio-visual-flamingo-hf
nvidia
5.0
OpenGVLab/InternVideo2_5_Chat_8B
OpenGVLab
5.0
OpenGVLab/VideoChat-Flash-Qwen2-7B_res448
OpenGVLab
5.0
OpenMOSS-Team/MOSS-VL-Instruct-0408
OpenMOSS-Team
5.0
OpenMOSS-Team/MOSS-VL-Realtime
OpenMOSS-Team
5.0
smdesai/SmolVLM2-2.2B-Instruct-4bit
smdesai
5.0
TencentARC/TimeLens-8B
TencentARC
5.0
Vision-CAIR/LongVU_Qwen2_7B
Vision-CAIR
5.0
yanziang/InternVideo3-8B-Instruct
yanziang
5.0
yaolily/TimeChat-Captioner-GRPO-7B
yaolily
5.0

78 models · click a column to sort

About the capability score: a 0–100 figure SyncDev calculates from each vendor's published specifications — context window, reasoning support, input modalities, tool calling, maximum output and how recently the model shipped. It measures breadth of capability, not benchmark performance, so a higher-scoring model is not automatically the better choice for your task.

Video-text-to-text model questions

It groups source-backed directory records for research and comparison. Review each model page and its primary documentation because collection membership alone does not prove quality, availability, or suitability.