Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
warning
This model is deprecated by newer models.
Embeddings
Apache 2.0 License
open_in_new Release Post

jina-embeddings-v2-base-de

German-English bilingual embeddings with SOTA performance
License
Apache-2.0
Release Date
calendar_month
2024-01-15
Input
abc
Text
arrow_forward
Output
more_horiz
Vector
Late Chunking help_outline
check_circle
Yes
Model Details
Parameters: 161M
Input Token Length: 8K
Output Dimension: 768
Base Model help_outline
jina-bert-v2-base-de
Trained Languages help_outline
2 languages
Related Models
link
jina-embeddings-v2-base-en
Available via
Jina API
AWS SageMaker
Microsoft Azure
Hugging Face
Air-gapped
I/O graph

Text

jina-embeddings-v2-base-de

Vector

Pareto fronthelp_outline
MTEB English · retrieval
MTEB English · sts
chevron_leftchevron_right
30M100M300M1B3.0B10B204060all-MiniLM-L6-v2e5-base-v2e5-mistral-7b-instructEmbeddingGemma-300MEVA-CLIP ViT-B/16gtr-t5-basegtr-t5-xlgtr-t5-xxljina-clip-v1jina-clip-v2jina-embedding-s-en-v1jina-embeddings-v2-smal…jina-embeddings-v3jina-embeddings-v4jina-embeddings-v5-text…jina-embeddings-v5-text…LongCLIP ViT-B/16nllb-siglip-largeOpenAI CLIP ViT-B/16Qwen3-Embedding-0.6BQwen3-Embedding-4Bsentence-t5-basesentence-t5-largesentence-t5-xlsentence-t5-xxljina-embeddings-v2-base…Parameters (log)nDCG@10
This model
On the front
Jina AI
Other
MTEB English · retrieval
44.10
Parameters
161M
Rank by score
27 / 39
Pareto front
Behind it
Value distributionhelp_outline
AUC 0.8452
Corpus
Translation pairs
Doc retrieval
0.6900.000.200.400.600.80
Related16.8%
Hard negative1.7%
Unrelated0.9%
Recommended cutoffs
FPR 0.1 · 0.499
FPR 0.01 · 0.690
FPR 0.001 · 0.796
FPR 0.0001 · 0.897
balanced · 0.368
AUC
0.8452
Noise ceiling
0.793
Recall cliff
0.195
Pairs measured
119 / 11k
Vector componentshelp_outline
-0.19-0.010.17
σ 0.0361 · 183k values
Embedding geometryhelp_outline
0768
Per-dimension mean, hover for a range
Noise floor
0.316
Effective dims
49 / 768
Language pairshelp_outline
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.158
Choose models to compare
Publications (1)
arXiv
February 26, 2024
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings

Overview

jina-embeddings-v2-base-de is a 161M-parameter bilingual text embedding model covering German and English with an 8,192-token context window. It maps semantically equivalent content in both languages into the same 768-dimensional embedding space, enabling cross-lingual retrieval without translation. The model was one of the first open-source bilingual embedding models to combine long-context support with balanced performance across both languages.

Methods

Built on a BERT-based backbone with symmetric bidirectional ALiBi positional encodings, the model processes both German and English through a unified 161M-parameter architecture producing 768-dimensional embeddings. Training included three stages: (1) multilingual pretraining on German-English parallel corpora, (2) contrastive fine-tuning on curated sentence pairs with hard negatives, and (3) cross-lingual alignment training to ensure semantically equivalent texts in German and English map to nearby regions of the embedding space. A key design choice was the bias-minimization objective, which counteracts the tendency of multilingual models to favor English grammatical structures — a documented failure mode in earlier multilingual embeddings. The 8,192-token window via ALiBi enables processing of full documents in either language without truncation.

Performance

The model outperformed Microsoft's E5-base while being less than a third of its size, and matched E5-large performance despite being 7× smaller. On WikiCLIR (English-to-German retrieval), STS17/STS22 (bidirectional semantic similarity), and BUCC (bilingual text alignment), it consistently outperformed models of comparable or larger size. The 322MB footprint enabled deployment on standard hardware. In 2026, jina-embeddings-v5-text-small supersedes this model for most applications, offering 32K context, 89 languages, and task-specific LoRA adapters. The v2-base-de model remains useful for German-English bilingual pipelines where the 8K context is sufficient.

Best Practice

Optimal for German-English bilingual retrieval: product search, support documentation, and content management where queries and documents may be in different languages. For documents exceeding 8,192 tokens, use semantic chunking or the `late_chunking parameter via the Jina API. The model integrates with Qdrant, Weaviate, MongoDB, and Milvus. For new multilingual projects spanning more than two languages, prefer jina-embeddings-v5-text-small` (89 languages, 32K context, LoRA adapters). CUDA-capable GPU recommended for production throughput.

Blogs that mention this model
September 27, 2024 • 15 minutes read
Migration From Jina Embeddings v2 to v3
We collected some tips to help you migrate from Jina Embeddings v2 to v3.
Alex C-G
Scott Martens
A digital upgrade theme with "V3" and a white "2", set against a green and black binary code background, with "Upgrade" centr
May 15, 2024 • 11 minutes read
Binary Embeddings: All the AI, 3.125% of the Fat
32-bits is a lot of precision for something as robust and inexact as an AI model. So we got rid of 31 of them! Binary embeddings are smaller, faster and highly performant.
Sofia Vasileva
Scott Martens
Futuristic digital 3D model of a coffee grinder with blue neon lights on a black background, featuring numerical data.
April 29, 2024 • 7 minutes read
Jina Embeddings and Reranker on Azure: Scalable Business-Ready AI Solutions
Jina Embeddings and Rerankers are now available on Azure Marketplace. Enterprises that prioritize privacy and security can now easily integrate Jina AI's state-of-the-art models right in their existing Azure ecosystem.
Susana Guzmán
Futuristic black background with a purple 3D grid, featuring the "Embeddings" and "Reranker" logos with a stylized "A".
January 31, 2024 • 16 minutes read
A Deep Dive into Tokenization
Tokenization, in LLMs, means chopping input texts up into smaller parts for processing. So why are embeddings billed by the token?
Scott Martens
Colorful speckled grid pattern with a mix of small multicolored dots on a black background, creating a mosaic effect.
January 26, 2024 • 13 minutes read
Jina Embeddings v2 Bilingual Models Are Now Open-Source On Hugging Face
Jina AI's open-source bilingual embedding models for German-English and Chinese-English are now on Hugging Face. We’re going to walk through installation and cross-language retrieval.
Scott Martens
Colorful "EMBEDDINGS" text above a pile of yellow smileys on a black background with decorative lines at the top.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.