I/O graph 1

multiple

Vector

Query

jina-colbert-v1-en

I/O graph 2

multiple

Vector

Document

jina-colbert-v1-en

Pareto front
30M100M7075808590ColBERTv2gte-reranker-modernbert…jina-reranker-v1-base-enjina-reranker-v1-tiny-enjina-reranker-v1-turbo-…jina-colbert-v1-enParameters (log)nDCG@10
This model
On the front
Jina AI
Other
LoCo
83.70
Parameters
137M
Rank by score
3 / 6
Pareto front
Behind it
Value distribution
AUC 0.5963
Corpus
Translation pairs
15.7801020
Related17.0%
Hard negative9.3%
Unrelated1.0%
Recommended cutoffs
FPR 0.1 · 9.27
FPR 0.01 · 15.78
FPR 0.001 · 21.53
FPR 0.0001 · 25.48
balanced · 7.98
AUC
0.5963
Noise ceiling
21.44
Recall cliff
-0.71
Pairs measured
100 / 9,200
Score by rank
12345678910
Mean score at each rank position
Choose models to compare

Overview

jina-colbert-v1-en is a 137M-parameter English late-interaction retrieval model that produces token-level embeddings (128-dimensional per token) instead of a single document vector. This enables cross-encoder quality with bi-encoder efficiency: documents are indexed once, and relevance scoring happens at query time via max-pooling and summation across token pairs. It supports 8,192-token documents and was the first Jina model to bring ColBERT-style retrieval to production.

Methods

The model employs a late-interaction architecture based on an adapted ColBERT approach. Instead of comparing entire documents at once, it processes queries and documents independently until the final matching stage. The document encoder processes text up to 8,192 tokens; the query encoder creates precise token-level representations. Each token in both query and document receives its own 128-dimensional embedding vector, preserving fine-grained semantic information lost in single-vector models. The late-interaction mechanism computes relevance scores by max-pooling over query tokens and summing over document tokens, avoiding the expensive all-to-all attention of cross-encoders while capturing token-level matching signals. The architecture uses 137M parameters with BERT-based encoders and ALiBi positional encodings for the 8,192-token context.

Performance

On the BEIR dataset collection, the model achieves superior performance: 49.4% on Arguana (vs. 46.5% for ColBERTv2), 79.5% on FEVER (vs. 78.8%), and 75.0% on TREC-COVID (vs. 72.6%). Most impressively, it shows a dramatic improvement on the LoCo benchmark for long-context understanding, scoring 83.7% compared to ColBERTv2's 74.3% — a 9.4-point gain that demonstrates the value of token-level representations for long documents. The model outperforms traditional single-vector embedding models while maintaining computational efficiency through the late-interaction approach. The 137M parameter count keeps it practical for production deployments.

Best Practice

Use this model when you need cross-encoder retrieval quality without cross-encoder latency. It requires a CUDA-capable GPU for optimal performance; CPU inference is possible for development. The 8,192-token document limit translates to approximately 6,000 words, suitable for most document types including academic papers and technical documentation. The model is English-only — for multilingual applications, use jina-colbert-v2. For production deployments, implement proper document chunking strategies and use vector similarity indexes (FAISS, Qdrant, Weaviate) for efficient retrieval. The model is particularly effective in RAG pipelines using frameworks like RAGatouille, which simplifies late-interaction implementation. For multilingual or higher-accuracy needs, consider jina-colbert-v2 or jina-embeddings-v4 (multi-vector mode).

Blogs that mention this model