Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
warning
This model is deprecated by newer models.
Embeddings
Apache 2.0 License
open_in_new Release Post

jina-embeddings-v2-base-es

Spanish-English bilingual embeddings with SOTA performance
License
Apache-2.0
Release Date
calendar_month
2024-02-14
Input
abc
Text
arrow_forward
Output
more_horiz
Vector
Late chunking help_outline
check_circle
Yes
Model Details
Parameters: 161M
Input Token Length: 8K
Output Dimension: 768
Base Model help_outline
jina-bert-v2-base-es
Trained Languages help_outline
2 languages
Related Models
link
jina-embeddings-v2-base-en
link
jina-embeddings-v2-base-de
link
jina-embeddings-v2-base-zh
Available via
Jina API
AWS SageMaker
Microsoft Azure
Hugging Face
Air-gapped
I/O graph

Text

jina-embeddings-v2-base-es

Vector

Pareto front help_outline
workspace_premium
MTEB English · retrieval
MTEB English · sts
chevron_leftchevron_right
30M100M300M1B3.0B657075808590bge-m3e5-base-v2EmbeddingGemma-300MEVA-CLIP ViT-B/16gtr-t5-basegtr-t5-largegtr-t5-xlgtr-t5-xxljina-clip-v2jina-embedding-b-en-v1jina-embedding-s-en-v1jina-embeddings-v3jina-embeddings-v4jina-embeddings-v5-text…jina-embeddings-v5-text…KaLM-mini-v2.5LongCLIP ViT-B/16multilingual-e5-basemultilingual-e5-large-i…nllb-siglip-largeOpenAI CLIP ViT-B/16Qwen3-Embedding-4Bsentence-t5-basesentence-t5-xlsentence-t5-xxljina-embeddings-v2-base…Parameters (log)Spearman
This model
On the front
Jina AI
Other
MTEB English · sts
83.50
Parameters
161M
Rank by score
11 / 39
Pareto front
On it
Value distributionhelp_outline
AUC 0.8383
Corpus
Translation pairs
Doc retrieval
0.6830.000.200.400.600.80
Related17.6%
Hard negative3.0%
Unrelated0.8%
Recommended cutoffs
FPR 0.1 · 0.509
FPR 0.01 · 0.683
FPR 0.001 · 0.775
FPR 0.0001 · 0.834
balanced · 0.365
AUC
0.8383
Noise ceiling
0.774
Recall cliff
0.163
Pairs measured
119 / 11k
Vector componentshelp_outline
-0.160.000.17
σ 0.0361 · 183k values
Embedding geometryhelp_outline
0768
Per-dimension mean, hover for a range
Noise floor
0.326
Effective dims
51 / 768
Language pairshelp_outline
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.315
Choose models to compare
Publications (1)
arXiv
February 26, 2024
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings

Overview

jina-embeddings-v2-base-es is a 161M-parameter bilingual text embedding model for Spanish and English, featuring an 8,192-token context window and 768-dimensional output. It was designed for cross-lingual retrieval in Spanish-speaking markets, mapping semantically equivalent Spanish and English content into a shared embedding space. The model addresses a critical gap: most embedding models at the time were English-centric, with Spanish performance significantly degraded.

Methods

The model uses a BERT-based backbone with symmetric bidirectional ALiBi positional encodings, 161M parameters, and a 768-dimensional output space. Training followed a three-stage process: (1) pretraining on Spanish-English parallel corpora, (2) contrastive fine-tuning with hard-negative mining on curated sentence pairs, and (3) cross-lingual alignment to ensure Spanish and English representations of the same concept cluster together. The ALiBi mechanism enables the 8,192-token context without learned positional embeddings, allowing the model to extrapolate beyond its 512-token training length. Mean pooling produces the final embedding vector.

Performance

The model outperformed significantly larger multilingual models (E5, BGE-M3) in Spanish-English retrieval tasks while being only 15–30% of their size. It demonstrated strong performance on MTEB Spanish subtasks, particularly in retrieval and clustering. The 8,192-token context window provided consistent performance on multi-page documents where most competing models required chunking. In 2026, jina-embeddings-v5-text-small supersedes this model, offering 89 languages, 32K context, and task-specific LoRA adapters. The v2-base-es model remains a cost-effective option for dedicated Spanish-English pipelines.

Best Practice

Ideal for Spanish-English bilingual search, content recommendation, and cross-lingual document analysis. For documents exceeding 8,192 tokens, use semantic chunking or the `late_chunking parameter via the Jina API. The model integrates with major vector databases and RAG frameworks. For new projects requiring multilingual coverage beyond Spanish-English, 32K context, or task-specific optimization, prefer jina-embeddings-v5-text-small`. CUDA-capable GPU recommended for production. Input text should be in Spanish or English; mixed-language input is supported but performance is optimized for the two trained languages.

Blogs that mention this model
April 29, 2024 • 7 minutes read
Jina Embeddings and Reranker on Azure: Scalable Business-Ready AI Solutions
Jina Embeddings and Rerankers are now available on Azure Marketplace. Enterprises that prioritize privacy and security can now easily integrate Jina AI's state-of-the-art models right in their existing Azure ecosystem.
Susana Guzmán
Futuristic black background with a purple 3D grid, featuring the "Embeddings" and "Reranker" logos with a stylized "A".
February 14, 2024 • 4 minutes read
Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.
Jina AI
Digital wireframe rendering of a Gothic-style cathedral, with colorful outlines and pointed spires on a dark background.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.