Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevancy.
Elastic Inference Service
Run Jina models natively inside Elasticsearch.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
warning
This model is deprecated by newer models.
Embeddings
Apache 2.0 License
open_in_new Release Post

jina-embeddings-v2-base-es

Spanish-English bilingual embeddings with SOTA performance
License
Apache-2.0
Release Date
calendar_month
2024-02-14
Input
abc
Text
arrow_forward
Output
more_horiz
Vector
Late Chunking help_outline
check_circle
Yes
Model Details
Parameters: 161M
Input Token Length: 8K
Output Dimension: 768
Base Model help_outline
open_in_new
jina-embeddings-v2-base-en
Trained Languages help_outline
2 languages
Related Models
link
jina-embeddings-v2-base-en
link
jina-embeddings-v2-base-de
link
jina-embeddings-v2-base-zh
Available via
Jina API
AWS SageMaker
Microsoft Azure
Hugging Face
Air-gapped
I/O graph

Text

jina-embeddings-v2-base-es

Vector

Pareto fronthelp_outline
MTEB English · retrieval
MTEB English · sts
chevron_leftchevron_right
30M100M300M1B3.0B10B204060all-MiniLM-L6-v2e5-base-v2e5-mistral-7b-instructEmbeddingGemma-300MEVA-CLIP ViT-B/16gtr-t5-basegtr-t5-xlgtr-t5-xxljina-clip-v1jina-clip-v2jina-embedding-s-en-v1jina-embeddings-v2-base…jina-embeddings-v2-smal…jina-embeddings-v3jina-embeddings-v4jina-embeddings-v5-text…jina-embeddings-v5-text…LongCLIP ViT-B/16nllb-siglip-largeOpenAI CLIP ViT-B/16Qwen3-Embedding-0.6BQwen3-Embedding-4Bsentence-t5-basesentence-t5-largesentence-t5-xlsentence-t5-xxljina-embeddings-v2-base…Parameters (log)nDCG@10
This model
On the front
Jina AI
Other
MTEB English · retrieval
46.40
Parameters
161M
Rank by score
24 / 39
Pareto front
Behind it
Value distributionhelp_outline
AUC 0.8383
Corpus
Translation pairs
Doc retrieval
0.6830.000.200.400.600.80
Related17.6%
Hard negative3.0%
Unrelated0.8%
Recommended cutoffs
FPR 0.1 · 0.509
FPR 0.01 · 0.683
FPR 0.001 · 0.775
FPR 0.0001 · 0.834
balanced · 0.365
AUC
0.8383
Noise ceiling
0.774
Recall cliff
0.163
Pairs measured
119 / 11k
Vector componentshelp_outline
-0.160.000.17
σ 0.0361 · 183k values
Embedding geometryhelp_outline
0768
Per-dimension mean, hover for a range
Noise floor
0.326
Effective dims
51 / 768
Language pairshelp_outline
de-ruen-deen-koen-zhja-ko
Cutoff spread across pairs: 0.315
Choose models to compare
Publications (1)
arXiv
February 26, 2024
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings

Overview

Jina Embeddings v2 Base Spanish is a groundbreaking bilingual text embedding model that addresses the critical challenge of cross-lingual information retrieval and analysis between Spanish and English content. Unlike traditional multilingual models that often show bias towards specific languages, this model delivers truly balanced performance across both Spanish and English, making it indispensable for organizations operating in Spanish-speaking markets or handling bilingual content. The model's most remarkable feature is its ability to generate geometrically aligned embeddings - when texts in Spanish and English express the same meaning, their vector representations naturally cluster together in the embedding space, enabling seamless cross-language search and analysis.

Methods

At the heart of this model lies an innovative architecture based on symmetric bidirectional ALiBi (Attention with Linear Biases), a sophisticated approach that enables processing of sequences up to 8,192 tokens without traditional positional embeddings. The model utilizes a modified BERT architecture with 161M parameters, incorporating Gated Linear Units (GLU) and specialized layer normalization techniques. Training follows a three-stage process: initial pre-training on a massive text corpus, followed by fine-tuning with carefully curated text pairs, and finally, hard-negative training to enhance discrimination between similar but semantically distinct content. This approach, combined with 768-dimensional embeddings, allows the model to capture nuanced semantic relationships while maintaining computational efficiency.

Performance

In comprehensive benchmark evaluations, the model demonstrates exceptional capabilities, particularly in cross-language retrieval tasks where it outperforms significantly larger multilingual models like E5 and BGE-M3 despite being only 15-30% of their size. The model achieves superior performance in retrieval and clustering tasks, showing particular strength in matching semantically equivalent content across languages. When tested on the MTEB benchmark, it exhibits robust performance across various tasks including classification, clustering, and semantic similarity. The extended context window of 8,192 tokens proves especially valuable for long-document processing, showing consistent performance even with documents spanning multiple pages - a capability most competing models lack.

Best Practice

To effectively utilize this model, organizations should ensure access to CUDA-capable GPU infrastructure for optimal performance. The model integrates seamlessly with major vector databases and RAG frameworks including MongoDB, Qdrant, Weaviate, and Haystack, making it readily deployable in production environments. It excels in applications such as bilingual document search, content recommendation systems, and cross-language document analysis. While the model shows impressive versatility, it's particularly optimized for Spanish-English bilingual scenarios and may not be the best choice for monolingual applications or scenarios involving other language pairs. For optimal results, input texts should be properly formatted in either Spanish or English, though the model handles mixed-language content effectively. The model supports fine-tuning for domain-specific applications, but this should be approached with careful consideration of the training data quality and distribution.
Blogs that mention this model
April 29, 2024 • 7 minutes read
Jina Embeddings and Reranker on Azure: Scalable Business-Ready AI Solutions
Jina Embeddings and Rerankers are now available on Azure Marketplace. Enterprises that prioritize privacy and security can now easily integrate Jina AI's state-of-the-art models right in their existing Azure ecosystem.
Susana Guzmán
Futuristic black background with a purple 3D grid, featuring the "Embeddings" and "Reranker" logos with a stylized "A".
February 14, 2024 • 4 minutes read
Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.
Jina AI
Digital wireframe rendering of a Gothic-style cathedral, with colorful outlines and pointed spires on a dark background.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.