Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
Data Ingestion
copyright CC BY-NC 4.0
open_in_new Release Post

jina-vlm

Multilingual vision-language model for visual question answering
License
copyright CC-BY-NC-4.0
Release Date
calendar_month
2025-12-04
Input
image
Image
abc
Text
arrow_forward
Output
abc
Text
Model Details
Parameters: 2.4B
Input Token Length: 32K
Input Image Size: 4096×4096
Base Model help_outline
open_in_new
Qwen3-1.7B-Base
open_in_new
SigLIP2 So400m
Trained Languages help_outline
39 languages
Supported Languages help_outline
93 languages
Apple Silicon Support help_outline
MLX
Related Models
link
jina-embeddings-v4
link
jina-reranker-m0
Available via
Jina API
Hugging Face
Air-gapped
I/O graph 1

Image

jina-vlm

Text

Text

I/O graph 2

Text

jina-vlm

Text

Pareto front help_outline
workspace_premium
DocVQA
AI2D
MMBench v1.1
MMMU
OCRBench
chevron_leftchevron_right
300M1B3.0B10B30B100B5060708090Aquila-VL-2BCambrian-1-34BCambrian-1-8BDeepSeek-VL2gemma-3-12b-itgemma-3-4b-itIdefics2-8BInternVL3-2BInternVL3-8BLLaVA-1.5-13BLLaVA-1.5-7BLLaVA-OneVision-0.5BLLaVA-OneVision-72BMiniCPM-V-2Molmo-72BMolmo-7B-DMolmoE-1Bmoondream2Ovis2-2BOvis2-8BPaliGemma-3B-mix-448Pixtral-12BQwen-VL-ChatQwen2.5-VL-72BQwen3-VL-2BSmolVLM-256MSmolVLM-500MSmolVLM-InstructSmolVLM2-2.2Bjina-vlmParameters (log)accuracy
This model
On the front
Jina AI
Other
AI2D
82.00
Parameters
2.5B
Rank by score
13 / 47
Pareto front
Behind it
Choose models to compare
Publications (2)
arXiv
December 29, 2025
Vision Encoders in Vision-Language Models: A Survey
ICLR 2026
December 04, 2025
Jina-VLM: Small Multilingual Vision Language Model

Overview

jina-vlm is a 2.4B-parameter multilingual vision-language model achieving state-of-the-art VQA performance among open 2B-scale VLMs. It couples a SigLIP2 vision encoder with a Qwen3 language decoder through an attention-pooling connector that enables token-efficient processing of arbitrary-resolution images. The model supports 29 languages and a 32K token context, making it suitable for document understanding, chart analysis, OCR, and multilingual visual question answering.

Methods

The architecture combines a SigLIP2 vision encoder with a Qwen3 language decoder, connected by an attention-pooling mechanism that enables token-efficient processing of arbitrary-resolution images. The key innovation is the attention-pooling connector, which applies 2×2 pooling to reduce 729 visual tokens per tile to 182 (a 4× token reduction), yielding 3.9× reduction in LLM prefill FLOPs and 4× reduction in KV-cache memory with minimal impact on benchmark scores. Images are processed via tiling: arbitrary-resolution images are divided into up to 12 tiles plus a thumbnail, each processed independently and then combined. The model was trained on a diverse multilingual VQA dataset covering 29 languages (Arabic, Chinese, English, Portuguese, Russian, Turkish, and others). A leave-one-out data mixture ablation study diagnosed which data categories (task, domain, modality, language) are necessary versus redundant, guiding efficient data allocation.

Performance

The model achieves the highest average score (72.3) across eight VQA benchmarks among 2B-scale VLMs: MathVista (59.4), AI2D (80.8), ChartQA (79.5), DocVQA (90.6), InfoVQA (65.9), RealWorldQA (64.9), OCRBench (778/1000), and MME (1582). It leads on multilingual multimodal understanding: MMMB (78.8) and Multilingual MMBench (74.3), covering Arabic, Chinese, English, Portuguese, Russian, and Turkish. OCR performance is strong at 778 on OCRBench (0–1000 scale). Text-only performance is competitive: MMLU (54.7), HellaSwag (75.6), though degraded on MMLU-Pro (30.3 vs. 46.4 base) due to vision-language integration. The 4× token reduction from attention pooling yields 3.9× reduction in LLM prefill FLOPs and 4× reduction in KV-cache memory with minimal benchmark impact.

Best Practice

The model is available on Hugging Face under CC-BY-NC-4.0 with weights and inference code. Supports images of arbitrary resolution through automatic tiling (up to 12 tiles plus thumbnail). Use thinking mode by enabling do_sample=True and temperature > 0 for complex reasoning tasks. The model handles 32K context length for extended conversations. For multilingual VQA, it supports 29 languages including English, Chinese, Arabic, German, Spanish, French, Italian, Japanese, Korean, Portuguese, Russian, Turkish, Vietnamese, Thai, Indonesian, Hindi, and Bengali. Best suited for document understanding, chart/diagram analysis, OCR tasks, and multilingual visual question answering. Limitations: counting tasks and fine-grained spatial reasoning may be affected by the tiling approach. For optimal inference, use bfloat16 precision on CUDA-capable GPUs. MLX inference is supported for Apple Silicon. For embedding-based retrieval over visual documents, pair with jina-embeddings-v4 or jina-reranker-m0.

Blogs that mention this model
December 04, 2025 • 7 minutes read
Jina-VLM: Small Multilingual Vision Language Model
New 2B vision language model achieves SOTA on multilingual VQA, no catastrophic forgetting on text-only tasks.
Jina AI
Artistic representation of "Vln" in vibrant, rainbow-like colors on a minimalistic white background, with a focus on color di
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.