Diankun
Diankun/Spatial-MLLM-v1.1-Instruct-135K
Diankun/Spatial-MLLM-v1.1-Instruct-135K is a source-linked Hugging Face repository assigned to Video-text-to-text. Its recorded task interface accepts video, text and produces text.
- Context
- —
- Max output
- —
- Input / 1M
- —
- Output / 1M
- —
- Output speed
- —
Accepts
- Video input
- Text input
Produces
- Text output