I/O graph

Text

jina-embeddings-v2-base-zh

Vector

Pareto front
30M100M300M1B3.0B10B506070acge_text_embeddingbge-base-zhbge-base-zh-v1.5bge-large-zh-noinstructbge-small-zhbge-small-zh-v1.5Conan-embedding-v1Dmeta-embedding-zh-smalle5-mistral-7b-instructgte-base-zhgte-large-zhgte-multilingual-basegte-Qwen1.5-7B-instructgte-Qwen2-1.5B-instructgte-Qwen2-7B-instructgte-small-zhluotuo-bert-mediumm3e-basem3e-largemultilingual-e5-basemultilingual-e5-largemultilingual-e5-smallpiccolo-base-zhpiccolo-large-zh-v2stella-base-zh-v2stella-base-zh-v3-1792dstella-large-zh-v2stella-mrl-large-zh-v3.…text2vec-base-chinesetext2vec-large-chinesexiaobu-embeddingxiaobu-embedding-v2Yinkazpoint_large_embedding_…Parameters (log)score
This model
On the front
Jina AI
Other
C-MTEB
63.79
Parameters
161M
Rank by score
23 / 40
Pareto front
Behind it
Value distribution
AUC 0.8373
Corpus
Translation pairs
Doc retrieval
0.6820.000.200.400.600.80
Related12.6%
Hard negative1.8%
Unrelated1.2%
Recommended cutoffs
FPR 0.1 · 0.475
FPR 0.01 · 0.682
FPR 0.001 · 0.796
FPR 0.0001 · 0.835
balanced · 0.353
AUC
0.8373
Noise ceiling
0.793
Recall cliff
0.113
Pairs measured
119 / 11k
Vector components
-0.180.000.18
σ 0.0361 · 183k values
Embedding geometry
0768
Per-dimension mean, hover for a range
Noise floor
0.308
Effective dims
56 / 768
Choose models to compare
Publications (1)

Overview

jina-embeddings-v2-base-zh is a 161M-parameter bilingual text embedding model for Chinese and English with an 8,192-token context window and 768-dimensional output. It was the first open-source model to seamlessly handle both Chinese and English with long-context support, addressing the unique tokenization challenges of Chinese character-based text in transformer architectures.

Methods

The model uses a BERT-based backbone with symmetric bidirectional ALiBi positional encodings, 161M parameters, and a 768-dimensional output space. Training followed a three-phase approach: initial pretraining on high-quality Chinese-English bilingual data, followed by primary and secondary fine-tuning stages with contrastive loss and hard-negative mining. The ALiBi mechanism enables the 8,192-token context window without learned positional embeddings. A notable improvement in the final release was a refined similarity score distribution that addressed score inflation issues present in the preview version, producing more discriminative and well-calibrated similarity scores.

Performance

On the C-MTEB (Chinese MTEB) leaderboard, the model demonstrated exceptional performance among models under 0.5GB, particularly excelling in Chinese-language tasks. It significantly outperformed OpenAI's text-embedding-ada-002 on Chinese-specific retrieval and similarity tasks while maintaining competitive performance on English tasks. The refined similarity score distribution improved discrimination between related and unrelated content in both languages. In 2026, jina-embeddings-v5-text-small supersedes this model with 89 languages, 32K context, and task-specific LoRA adapters.

Best Practice

Optimal for Chinese-English bilingual retrieval, cross-lingual document search, and multilingual content analysis. For documents exceeding 8,192 tokens, use semantic chunking or the `late_chunking parameter via the Jina API. The model integrates with major vector databases and RAG frameworks. For new multilingual projects, prefer jina-embeddings-v5-text-small` (89 languages, 32K context, LoRA adapters). CUDA-capable GPU recommended for production throughput. Input text should be in Chinese or English; the model handles both languages natively without translation.

Blogs that mention this model