Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
warning
This model is deprecated by newer models.
Embeddings
Apache 2.0 License

jina-embedding-b-en-v1

The first version of the Jina Embedding model, the OG.
License
Apache-2.0
Release Date
calendar_month
2023-06-17
Input
abc
Text
arrow_forward
Output
more_horiz
Vector
Model Details
Parameters: 110M
Input Token Length: 512
Output Dimension: 768
Base Model help_outline
open_in_new
T5-Base Encoder
Trained Languages help_outline
1 languages
Related Models
link
jina-embeddings-v2-base-en
link
jina-embeddings-v3
Available via
Hugging Face
Air-gapped
I/O graph

Text

jina-embedding-b-en-v1

Vector

Pareto fronthelp_outline
MTEB English · retrieval
MTEB English · sts
chevron_leftchevron_right
30M100M300M1B3.0B10B204060e5-base-v2e5-mistral-7b-instructEmbeddingGemma-300MEVA-CLIP ViT-B/16gtr-t5-basegtr-t5-xlgtr-t5-xxljina-clip-v1jina-clip-v2jina-embedding-l-en-v1jina-embedding-s-en-v1jina-embeddings-v2-base…jina-embeddings-v2-smal…jina-embeddings-v3jina-embeddings-v4jina-embeddings-v5-text…jina-embeddings-v5-text…LongCLIP ViT-B/16nllb-siglip-largeOpenAI CLIP ViT-B/16Qwen3-Embedding-0.6BQwen3-Embedding-4Bsentence-t5-basesentence-t5-largesentence-t5-xlsentence-t5-xxljina-embedding-b-en-v1Parameters (log)nDCG@10
This model
On the front
Jina AI
Other
MTEB English · retrieval
44.03
Parameters
110M
Rank by score
28 / 39
Pareto front
Behind it
Choose models to compare
Publications (1)
EMNLP 2023
July 20, 2023
Jina Embeddings: A Novel Set of High-Performance Sentence Embedding Models

Overview

jina-embedding-b-en-v1 was Jina AI's first publicly released text embedding model: a 110M-parameter bidirectional encoder that maps English text into 768-dimensional vectors. Designed for semantic search, similarity comparison, and content recommendation at a time when 512-token context was the industry standard, it established Jina's entry into the open-source embedding space. The model is now legacy and has been superseded by the v2 and v3 families, which support 8K+ token contexts and multilingual input.

Methods

The model uses a T5-encoder architecture with mean pooling to produce fixed-length 768-dimensional representations. Training followed a two-phase contrastive recipe on the Linnaeus-Clean dataset (385M sentence pairs filtered from 1.6B candidates): first, InfoNCE loss on positive text pairs to learn semantic alignment; second, triplet loss to sharpen discrimination between similar but semantically distinct texts. A key design decision was the mean-pooling strategy, which averages token-level representations to produce a single vector capturing global sentence meaning. The 512-token context window was standard for its era but limits the model's usefulness for long-document applications.

Performance

On STS12, the model achieved a correlation score of 0.751, outperforming all-mpnet-base-v2 and all-minilm-l6-v2 at release. It demonstrated strong performance across standard English sentence-embedding tasks while maintaining fast inference times suitable for production. Its 512-token context window and English-only focus made it unsuitable for the long-document and multilingual workloads that became standard by 2025. The model has been superseded by jina-embeddings-v2-base-en (8K context, bidirectional ALiBi) and jina-embeddings-v3 (570M params, 89 languages, task-specific LoRA adapters).

Best Practice

This model is legacy and should not be used for new projects. If migrating from jina-embedding-b-en-v1, upgrade to jina-embeddings-v5-text-small for current workloads — it supports 32K token context, 89 languages, and task-specific LoRA adapters. For code-specific embedding needs, use jina-code-embeddings-1.5b. The model requires CUDA-capable hardware for optimal performance and accepts inputs up to 512 tokens. It is not suitable for multilingual content, long documents, or code-centric applications.

Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.