I/O graph 1

Text

jina-embeddings-v5-omni-small

Image

Task

Vector

I/O graph 2

Text

jina-embeddings-v5-omni-small

Audio

Task

Vector

I/O graph 3

Text

jina-embeddings-v5-omni-small

Video

Task

Vector

I/O graph 4

multiple

Vector

Text

jina-embeddings-v5-omni-small

PDF

Task

Pareto front
300M1B3.0B10B20304050BidirLM-Omni-2.5B-Embed…e5-omni-3Be5-omni-7Bjina-embeddings-v5-omni…LanguageBindlarger_clap_generallarger_clap_musicLCO-Embedding-Omni-3Bmsclap-2022MuQ-MuLan-largeOmni-Embed-Nemotron-3BQwen2-Audio-7BQwen2.5-Omni-3Bspeecht5_multimodalwav2clipjina-embeddings-v5-omni…Parameters (log)score
This model
On the front
Jina AI
Other
MAEB
49.96
Parameters
1.6B
Rank by score
5 / 21
Pareto front
On it
Value distribution
AUC 0.8269
Corpus
Translation pairs
Doc retrieval
Code
Image / banner
Image / logo
Task
classification
clustering
retrieval.passage
retrieval.query
retrieval.query → retrieval.passage
text-matching
0.8080.400.500.600.700.800.90
Related10.9%
Hard negative2.1%
Unrelated0.9%
Recommended cutoffs
FPR 0.1 · 0.717
FPR 0.01 · 0.808
FPR 0.001 · 0.860
FPR 0.0001 · 0.878
balanced · 0.697
AUC
0.8269
Noise ceiling
0.855
Recall cliff
0.566
Pairs measured
119 / 11k
This model shares its text tower with jina-embeddings-v5-text-small. The distributions here are that model's, which it matches to fp16 wire precision.
Vector components
-0.160.010.18
σ 0.0313 · 244k values
Embedding geometry
01024
Per-dimension mean, hover for a range
Noise floor
0.300
Effective dims
70 / 1024
Dimension truncation
32641282565121024
text-matching · Cutoff by requested dimensions
Language pairs
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.024
Choose models to compare
Publications (1)

Overview

jina-embeddings-v5-omni-small is a ~1.74B-parameter multimodal embedding model that accepts text, images, video, audio, and PDF, producing embeddings in a shared 1024-dimensional vector space. It is the default embedding model for Jina's API and offers the best accuracy-efficiency tradeoff among Jina's multimodal embedding models. Text-only outputs are bit-identical to jina-embeddings-v5-text-small, enabling seamless migration from text-only to multimodal without reindexing.

Methods

The model is trained in a third stage extending jina-embeddings-v5-text-small. The Qwen3-0.6B-Base text backbone and all four task-specific LoRA adapters are frozen during multimodal training; only the cross-modal projectors are newly trained. A SigLIP2 Large vision encoder handles images and video (32 uniformly sampled frames per video). A Whisper-large-v3 audio encoder handles audio input. PDF pages are rendered as images and processed through the vision pathway. Training uses contrastive loss with cross-modal hard negatives to align visual and audio representations with the existing text embedding space. The 1024-dimensional output supports Matryoshka truncation down to 32 dimensions. The 32K token context window covers text input. The model supports quantized and MLX inference for Apple Silicon.

Performance

Text-only performance is bit-identical to jina-embeddings-v5-text-small (67.0 MMTEB task-level average, 71.7 English MTEB) — the text backbone and LoRA adapters are untouched during multimodal training. On cross-modal retrieval, the model demonstrates strong alignment across text-image, text-audio, and text-video tasks. PDF page retrieval is handled through the vision pathway. The omni-small model offers the best accuracy-efficiency tradeoff among Jina multimodal embedding models for server deployment, with 1024-dimensional embeddings providing more discriminative power than the 768-dimensional omni-nano variant.

Best Practice

Select the appropriate LoRA adapter: retrieval, text-matching, clustering, or classification. For multimodal inputs via the API, pass image URLs, audio file URLs, video file URLs, or PDF URLs directly — the model routes each modality through the appropriate encoder. Supported audio formats include WAV, MP3, FLAC, OGG, M4A, and Opus. Video inputs are processed as 32 uniformly sampled frames. Mix modalities freely within a single batch: the embedding space is shared across all modalities. Use cosine similarity for comparison. Matryoshka truncation from 1024 to 32 dimensions is supported. Text-only embeddings are drop-in compatible with jina-embeddings-v5-text-small — no reindexing needed when upgrading. For edge deployments without GPU, use jina-embeddings-v5-omni-nano instead.

Blogs that mention this model