Elastic
Jina AI
News
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevancy.
Elastic Inference Service
Run Jina models natively inside Elasticsearch.
MCP terminalCLIarticlellms.txtsmart_toyAgentsdata_objectSchemamenu_bookDocs



Log in
login
Embeddings
copyright CC BY-NC 4.0
open_in_new Release Post

jina-clip-v2

Multilingual multimodal embeddings for texts and images
License
copyright CC-BY-NC-4.0
Release Date
calendar_month
2024-11-05
Input
image
Image
abc
Text
arrow_forward
Output
more_horiz
Vector
Matryoshka Dimensions help_outline
64
128
256
512
768
1024
Model Details
Parameters: 865M
Input Token Length: 8K
Input Image Size: 512×512
Output Dimension: 1024
Base Model help_outline
open_in_new
XLM-RoBERTa Large
Trained Languages help_outline
32 languages
Supported Languages help_outline
108 languages
Related Models
link
jina-clip-v1
Available via
Elastic Inference Service
Jina API
AWS SageMaker
Microsoft Azure
Google Cloud
Hugging Face
Air-gapped
I/O graph 1

Text

jina-clip-v2

Vector

I/O graph 2

Image

jina-clip-v2

Vector

Value distributionhelp_outline
AUC 0.8383
Corpus
Translation pairs
Doc retrieval
Image / banner
Image / logo
Task
default
retrieval.query
0.6040.000.200.400.60
Related20.2%
Hard negative3.6%
Unrelated0.8%
Hover or click the chart to move the cutoff
Recommended cutoffs
FPR 0.1 · 0.508
FPR 0.01 · 0.604
FPR 0.001 · 0.668
FPR 0.0001 · 0.714
balanced · 0.403
AUC
0.8383
Noise ceiling
0.666
Recall cliff
0.224
Pairs measured
119 / 11k
Vector componentshelp_outline
-0.43-0.070.29
σ 0.0313 · 244k values
Embedding geometryhelp_outline
01024
Per-dimension mean, hover for a range
Noise floor
0.420
Effective dims
52 / 1024
Language pairshelp_outline
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.064
Choose models to compare
Publications (2)
ICLR 2026
January 22, 2026
Embedding Compression via Spherical Coordinates
ICLR 2025
December 12, 2024
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Overview

Jina CLIP v2 revolutionizes multimodal AI by bridging the gap between visual and textual understanding across 89 languages. This model solves critical challenges in global e-commerce, content management, and cross-cultural communication by enabling accurate image-text matching regardless of language barriers. For businesses expanding internationally or managing multilingual content, it eliminates the need for separate models per language or complex translation pipelines. The model particularly shines in scenarios requiring precise visual search across language boundaries, such as global marketplace product discovery or multilingual digital asset management.

Methods

At its core, Jina CLIP v2 employs a sophisticated dual-encoder architecture that combines a Jina XLM-RoBERTa text encoder (561M parameters) with an EVA02-L14 vision encoder (304M parameters). The text encoder processes content in 89 languages with a massive context window of 696,320 tokens, while the vision encoder handles high-resolution images up to 512x512 pixels. The model introduces innovative Matryoshka representation learning, which enables dynamic embedding dimension adjustment from 1024 down to 64 dimensions while preserving performance. This architecture processes both text and images through their respective encoders, projecting them into a shared semantic space where similar concepts align regardless of their original modality or language.

Performance

The model achieves state-of-the-art performance with 98.0% accuracy on Flickr30k image-to-text retrieval tasks, surpassing both its predecessor and NLLB-CLIP-SigLIP. In multilingual scenarios, it demonstrates up to 4% improvement over NLLB-CLIP-SigLIP in cross-lingual image retrieval tasks, despite having fewer parameters than its largest competitor. The model maintains strong performance even when embeddings are compressed - reducing dimensions by 75% still preserves over 99% of performance across text, image, and cross-modal tasks. On the comprehensive Multilingual MTEB benchmarks, it achieves 69.86% on retrieval and 67.77% on semantic similarity tasks, performing competitively with specialized text embedding models.

Best Practice

For optimal deployment, users should consider several key factors. The model requires CUDA-capable hardware for efficient processing, with memory requirements scaling based on batch size and image resolution. To optimize API costs and performance, resize images to 512x512 pixels before processing - larger images are automatically tiled, increasing token usage and processing time. The model excels at matching images with descriptive text across languages but may struggle with abstract concepts or highly specialized domain-specific content. It's particularly effective for e-commerce product search, content recommendation systems, and visual search applications, but may not be suitable for tasks requiring fine-grained visual detail analysis or highly specialized domain expertise. When using the Matryoshka representation feature, consider the trade-off between dimension reduction and performance - while 64-dimension embeddings maintain strong performance, critical applications may benefit from higher dimensions.
Blogs that mention this model
November 21, 2024 • 9 minutes read
Jina CLIP v2: Multilingual Multimodal Embeddings for Text and Images
Jina-CLIP v2, a 0.9B multimodal embedding model with multilingual support of 89 languages, high image resolution at 512x512, and Matryoshka representations.
Jina AI
Digital number "2" displayed in a mosaic of colorful squares against a dark background, creating a futuristic vibe.
October 29, 2024 • 11 minutes read
Beyond CLIP: How Jina-CLIP Advances Multimodal Search
Learn how Jina-CLIP enhances OpenAI's CLIP with better retrieval accuracy and more diverse results through unified text-image embeddings.
Bo Wang
Alex C-G
Abstract digital landscape with wave-like green and pink dunes against a dark background, conveying a tranquil atmosphere.
June 05, 2024 • 9 minutes read
Jina CLIP v1: A Truly Multimodal Embeddings Model for Text and Image
Jina AI's new multimodal embedding model not only outperforms OpenAI CLIP in text-image retrieval, it's a solid image embedding model and state-of-the-art text embedding model at the same time. You don't need different models for different modalities any more.
Sofia Vasileva
Scott Martens
Susana Guzmán
Abstract 3D render of a neon blue and green grid pattern on a black background, creating a sense of depth.
September 27, 2024 • 15 minutes read
Migration From Jina Embeddings v2 to v3
We collected some tips to help you migrate from Jina Embeddings v2 to v3.
Alex C-G
Scott Martens
A digital upgrade theme with "V3" and a white "2", set against a green and black binary code background, with "Upgrade" centr
August 30, 2024 • 10 minutes read
Jina ColBERT v2: Multilingual Late Interaction Retriever for Embedding and Reranking
Jina ColBERT v2 supports 89 languages with superior retrieval performance, user-controlled output dimensions, and 8192 token-length.
Jina AI
Dark-themed coding interface displaying English and Japanese characters with "JINA COLBERT V2" highlighted in the center.
Search Foundation
Reader
Embeddings
Reranker
Elastic Inference Service
open_in_new
Get Jina API key
Rate Limit
API Status
Terms
Security
Terms & Conditions
Privacy
Manage Cookies
Do Not Sell or Share My Personal Information
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
Elastic © 2020-2026.
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.