I/O graph 1

Text

jina-clip-v1

Vector

I/O graph 2

Image

jina-clip-v1

Vector

Pareto front
300M1B3.0B10B20406080100bge-visualized-m3BiModernVBERTDSE-Phi3dse-qwen2-2b-mrl-v1e5-vjina-clip-v2jina-embeddings-v3jina-embeddings-v4llama-nemoretriever-col…MoCa-3Bnemotron-colembed-vl-8b…nllb-siglip-largeONE-PEACESauerkrautLM-ColLFM2-45…SauerkrautLM-ColMinistr…siglip-so400m-patch14-3…SigLIP2-L-512/16VLM2VecVultronRetrieverFlash-Q…jina-clip-v1Parameters (log)nDCG@5
This model
On the front
Jina AI
Other
ViDoRe v1
17.72
Parameters
223M
Rank by score
32 / 33
Pareto front
On it
Value distribution
AUC 0.8440
Corpus
Translation pairs
Doc retrieval
Image / banner
0.6750.000.200.400.600.80
Related20.2%
Hard negative3.2%
Unrelated1.2%
Recommended cutoffs
FPR 0.1 · 0.478
FPR 0.01 · 0.675
FPR 0.001 · 0.785
FPR 0.0001 · 0.823
balanced · 0.376
AUC
0.8440
Noise ceiling
0.779
Recall cliff
0.123
Pairs measured
119 / 11k
Vector components
-0.18-0.010.15
σ 0.0361 · 183k values
Embedding geometry
0768
Per-dimension mean, hover for a range
Noise floor
0.283
Effective dims
55 / 768
Choose models to compare
Publications (1)

Overview

jina-clip-v1 is a 223M-parameter multimodal embedding model that achieves state-of-the-art performance on both text-to-text and text-to-image retrieval tasks — a first among CLIP-family models. Unlike standard CLIP models that sacrifice text-only retrieval for cross-modal alignment, jina-clip-v1 uses multi-task contrastive training to maintain strong performance across all retrieval combinations. It supports 12,288-token text context, 100× the original CLIP's 77-token limit.

Methods

The architecture combines an adapted Jina BERT v2 text encoder (supporting 12,288 tokens via ALiBi) with the EVA-02 image encoder from the Beijing Academy for AI. The key innovation is the multi-task contrastive training method: the model is trained simultaneously on (1) image-caption pairs for cross-modal alignment, (2) text-text pairs to preserve text-only retrieval quality, and (3) AI-generated longer text descriptions of images to improve robustness to varying text lengths. A final stage uses hard negative text triplets to sharpen semantic distinction. This multi-task approach prevents the catastrophic forgetting of text-only retrieval that plagues standard CLIP training. The image encoder processes 224×224 pixel tiles, with each tile consuming 1,000 tokens of processing capacity.

Performance

In text-only retrieval, the model achieves a 165% performance increase over OpenAI CLIP (0.429 vs. 0.162 on MTEB). For image tasks: 2% better in text-to-image retrieval (0.899), 6% in image-to-text retrieval (0.803), and 12% in image-to-image retrieval (0.916). On zero-shot visual classification (CIFAR-100), it outperforms standard CLIP. Cross-modal performance on Flickr8k/30k and MSCOCO Captions is competitive with specialized single-modality models. The 12,288-token text context is a major advantage for retrieving long-form text documents alongside images. In 2026, jina-clip-v2 supersedes this model with 89-language support and 512×512 image processing.

Best Practice

The model processes images in 224×224 pixel tiles; each tile consumes 1,000 tokens. A 750×500 pixel image requires 12 tiles (12,000 tokens). Text requires approximately 1.1 tokens per word. Plan token budgets accordingly for mixed-modal queries. The model currently supports English only — for multilingual applications, use jina-clip-v2. It is available through the Jina Embeddings API and as an open-source release on Hugging Face under the Apache 2.0 license. For production environments, AWS Marketplace and Azure deployment options provide optimized infrastructure. Use cosine similarity for cross-modal comparison. For new multimodal projects, prefer jina-clip-v2 (89 languages, 512×512 images, MRL support) or jina-embeddings-v4 (32K context, multi-vector mode, code support).

Blogs that mention this model