Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
Embeddings
copyright CC BY-NC 4.0
open_in_new Release Post

jina-clip-v2

Multilingual multimodal embeddings for texts and images
License
copyright CC-BY-NC-4.0
Release Date
calendar_month
2024-11-05
Input
image
Image
abc
Text
arrow_forward
Output
more_horiz
Vector
Matryoshka Dimensions help_outline
64
128
256
512
768
1024
Model Details
Parameters: 865M
Input Token Length: 8K
Input Image Size: 512×512
Output Dimension: 1024
Base Model help_outline
open_in_new
EVA02-L-14
link
jina-embeddings-v3
Trained Languages help_outline
32 languages
Supported Languages help_outline
108 languages
Related Models
link
jina-clip-v1
Available via
Elastic Inference Service
Jina API
AWS SageMaker
Microsoft Azure
Google Cloud
Hugging Face
Air-gapped
I/O graph 1

Text

jina-clip-v2

Vector

I/O graph 2

Image

jina-clip-v2

Vector

Pareto fronthelp_outline
ViDoRe v1
CLIP Benchmark
CLIP Benchmark · t2i-recall@5
MTEB English · retrieval
MTEB English · sts
chevron_leftchevron_right
100M300M1B3.0B707580859095CLIP-ViT-g-14-laion2BCLIPA-v2 ViT-bigG/14-336DFN2B-CLIP-ViT-B-16DFN5B-CLIP-ViT-H-14-378EVA02-CLIP-E-14MetaCLIP ViT-H/14MetaCLIP ViT-L/14NLLB-CLIP-baseNLLB-CLIP-largenllb-siglip-basenllb-siglip-largeOpenAI CLIP RN101OpenAI CLIP RN50x4OpenAI CLIP ViT-L/14siglip-base-patch16-224siglip-base-patch16-512siglip-large-patch16-384siglip-so400m-patch14-3…jina-clip-v2Parameters (log)i2t-recall@5
This model
On the front
Jina AI
Other
CLIP Benchmark
89.73
Parameters
865M
Rank by score
34 / 56
Pareto front
Behind it
Value distributionhelp_outline
AUC 0.8383
Corpus
Translation pairs
Doc retrieval
Image / banner
Image / logo
Task
default
retrieval.query
0.6040.000.200.400.60
Related20.2%
Hard negative3.6%
Unrelated0.8%
Recommended cutoffs
FPR 0.1 · 0.508
FPR 0.01 · 0.604
FPR 0.001 · 0.668
FPR 0.0001 · 0.714
balanced · 0.403
AUC
0.8383
Noise ceiling
0.666
Recall cliff
0.224
Pairs measured
119 / 11k
Vector componentshelp_outline
-0.43-0.070.29
σ 0.0313 · 244k values
Embedding geometryhelp_outline
01024
Per-dimension mean, hover for a range
Noise floor
0.420
Effective dims
52 / 1024
Language pairshelp_outline
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.064
Choose models to compare
Publications (2)
ICLR 2026
January 22, 2026
Embedding Compression via Spherical Coordinates
ICLR 2025
December 12, 2024
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Overview

jina-clip-v2 is an 865M-parameter multilingual multimodal embedding model supporting 89 languages and 512×512 image input. It extends jina-clip-v1 with cross-lingual image-text alignment, Matryoshka representation learning for dimension reduction (1024→64), and improved performance on visually rich documents. It achieves state-of-the-art cross-lingual image retrieval while maintaining strong text-only and single-modal performance.

Methods

The model employs a dual-encoder architecture combining a Jina XLM-RoBERTa text encoder (561M parameters, 89 languages, 696,320-token context) with an EVA02-L14 vision encoder (304M parameters, 512×512 pixel input). Training uses a multi-task, multi-stage contrastive learning paradigm: the model is trained on text pairs, text triplets, and image-text pairs simultaneously to support both text-only and cross-modal tasks. The training dataset was expanded to include multilingual texts from 29 non-English languages (including Hindi, Chinese, German, French) and images of visually rich documents. Matryoshka representation learning enables embedding dimension reduction from 1024 to 64 dimensions while preserving over 99% of performance. The model uses Last-Token-Pooling for the final embedding.

Performance

The model achieves 98.0% accuracy on Flickr30k image-to-text retrieval, surpassing both jina-clip-v1 and NLLB-CLIP-SigLIP. In multilingual scenarios, it demonstrates up to 4% improvement over NLLB-CLIP-SigLIP in cross-lingual image retrieval. Embedding compression is remarkably effective: reducing dimensions by 75% (1024→256) still preserves over 99% of performance across text, image, and cross-modal tasks. On Multilingual MTEB, it achieves 69.86% on retrieval and 67.77% on semantic similarity, performing competitively with specialized text embedding models. The 89-language support and visually-rich document handling make it uniquely suited for global e-commerce and content management.

Best Practice

Resize images to 512×512 pixels before processing — larger images are automatically tiled, increasing token usage and processing time. The model excels at matching images with descriptive text across 89 languages, making it ideal for global e-commerce product search, content recommendation, and multilingual visual search. It may struggle with abstract concepts or highly specialized domain content. When using Matryoshka dimension reduction, 64-dimension embeddings maintain strong performance for indexing; use 512+ dimensions for re-ranking critical applications. The model requires CUDA-capable hardware. For text-only workloads, jina-embeddings-v5-text-small offers better accuracy-per-parameter. For code retrieval, use jina-code-embeddings-1.5b. Available via Jina API, AWS, Azure, and GCP marketplaces.

Blogs that mention this model
July 31, 2025 • 12 minutes read
How Image Resolution Impacts Visual Document Retrieval
Image resolution is crucial for embedding visually rich documents. Too small and models miss key details; too large and they can't connect the parts.
Maximilian Werk
Michael Günther
Scott Martens
Abstract composition with a dark background featuring a flower-like design, radiant eye-like feature, rainbow-colored curved
July 25, 2025 • 8 minutes read
JinaVDR: New Visual Document Retrieval Benchmark with 95 Tasks in 20 Languages
JinaVDR is a new benchmark spanning 95 tasks across 20 languages for visual document retrieval, soon on MTEB.
Maximilian Werk
Alex C-G
Black-and-white design for "Jinavor Benchmark" with bold text. Below, "Visual Docs: 95 Tasks: 20 Languages" appears; an abstr
June 25, 2025 • 12 minutes read
Jina Embeddings v4: Universal Embeddings for Multimodal Multilingual Retrieval
Jina Embeddings v4 is a 3.8 billion parameter universal embedding model for multimodal and multilingual retrieval that supports both single-vector and multi-vector embedding outputs.
Jina AI
Word "Embeddings" followed by a numeric or symbol representation, displayed in multiple colors on a technology-themed, colorf
May 28, 2025 • 4 minutes read
Correlations: Vibe-Testing Embeddings in GUI
As serious as we are about MTEB, we also love vibe-testing. Correlations is a simple GUI we use for validating citations in DeepSearch, debugging late chunking, and vibe-testing embeddings. Now it's open-source.
Jina AI
Technical screen showing green and yellow visual data, including charts in the lower half and a heat-map-like visualization a
May 25, 2025 • 21 minutes read
What We Learned at ICLR2025
We collect some most interesting papers in ICLR 2025, featuring TIPS, FlexPrefill, Zero-Shot Rerankers, SVD-LLM, Hymba etc.
Jina AI
Three people smiling on a stage at a conference with an ICLR banner visible, suggesting a warm and lively event atmosphere.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.