Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevancy.
Elastic Inference Service
Run Jina models natively inside Elasticsearch.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login

Contact sales

Grow your business with Jina AI.

nights_stay Sales team away, back in 3h

Two Ways to Purchase

Subscribe to our API or purchase through cloud providers.
radio_button_unchecked
cloud
With 3 cloud service providers
Using AWS or Azure? You can deploy our models directly on your company's cloud platform and handle billing through the CSP account.
AWS SageMaker
Embeddings
Reranker
Microsoft Azure
Embeddings
Reranker
Google Cloud
Embeddings
radio_button_checked
With Jina Search Foundation API
The easiest way to access all of our products. Top-up tokens as you go.
Top up this API key with more tokens
Depending on your location, you may be charged in USD, EUR, or other currencies. Taxes may apply.
Please input the right API key to top up
Understand the rate limit
Rate limits are the maximum number of requests that can be made to an API within a minute per IP address/API key (RPM). Find out more about the rate limits for each product and tier below.
keyboard_arrow_down
Rate Limit
Rate limits are tracked in three ways: RPM (requests per minute), and TPM (tokens per minute). Limits are enforced per IP/API key and will be triggered when either the RPM or TPM threshold is reached first. When you provide an API key in the request header, we track rate limits by key rather than IP address.
ProductAPI EndpointDescriptionarrow_upwardw/o API Keykey_offw/ Free API Keykeyw/ Paid API Keykeyw/ Premium API KeykeyAverage LatencyToken Usage CountingAllowed Request
Reader APIhttps://r.jina.aiConvert URL to LLM-friendly text20 RPM500 RPM500 RPMtrending_up5000 RPM7.9sCount the number of tokens in the output response.GET/POST
Reader APIhttps://s.jina.aiSearch the web and convert results to LLM-friendly textblock100 RPM100 RPMtrending_up1000 RPM2.5sEvery request costs a fixed number of tokens, starting from 10000 tokensGET/POST
Embedding APIhttps://api.jina.ai/v1/embeddingsConvert text/images to fixed-length vectorsblock100 RPM & 100,000 TPM500 RPM & 2,000,000 TPMtrending_up5,000 RPM & 50,000,000 TPM
ssid_chart
depends on the input size
help
Count the number of tokens in the input request.POST
Reranker APIhttps://api.jina.ai/v1/rerankRank documents by queryblock100 RPM & 100,000 TPM500 RPM & 2,000,000 TPMtrending_up5,000 RPM & 50,000,000 TPM
ssid_chart
depends on the input size
help
Count the number of tokens in the input request.POST
Classifier APIhttps://api.jina.ai/v1/trainTrain a classifier using labeled examplesblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokens × num_itersPOST
Classifier API (Few-shot)https://api.jina.ai/v1/classifyClassify inputs using a trained few-shot classifierblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokensPOST
Classifier API (Zero-shot)https://api.jina.ai/v1/classifyClassify inputs using zero-shot classificationblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokens + label_tokensPOST
Segmenter APIhttps://api.jina.ai/v1/segmentTokenize and segment long text20 RPM200 RPM200 RPM1,000 RPM0.3sToken is not counted as usage.GET/POST
DeepSearchhttps://deepsearch.jina.ai/v1/chat/completionsReason, search and iterate to find the best answerblock50 RPM50 RPM500 RPM56.7sCount the total number of tokens in the whole process.POST

FAQ

Jina AI × Elastic

handshake
Will the Jina brand be preserved?
keyboard_arrow_down
Yes. Jina is a model brand within Elastic. Think of it like Qwen to Alibaba, GPT to OpenAI, or Kimi to Moonshot. The legal identity has moved to Elastic, which lets Jina AI focus purely on search foundation models as a brand.
handshake
What will Jina AI focus on going forward?
keyboard_arrow_down
Embeddings, rerankers, and small models for better search, including multimodal and reader models. Our mission isn't accomplished yet, and we've never been shy about the goal: to be a leading search model provider.
handshake
Will the API and cloud marketplace offerings continue?
keyboard_arrow_down
Yes. The Reader API, Embeddings API, and Reranker API continue to be developed and maintained, and models continue to be published to cloud marketplace platforms. You can use our API services as before. The only exception is that we cannot serve entities or countries subject to U.S. export controls.
handshake
Will you still release open-weights models on Hugging Face?
keyboard_arrow_down
Yes. At Elastic, Jina continues to push the frontier of search foundation models, and we keep releasing open-weights models.
handshake
Under which license will these open models be released?
keyboard_arrow_down
We continue releasing under CC-BY-NC 4.0 unless circumstances change, which is unlikely. Legacy v1 and v2 generation models remain Apache-2.0. CC-BY-NC 4.0 permits non-commercial use only: the weights are free to download, evaluate, benchmark, and use in research, but running them in a commercial product requires a commercial license. Since August 10, 2026, Elastic sells that license separately from any Elasticsearch subscription, so contact Elastic Sales to arrange one.
handshake
Will you continue publishing research papers?
keyboard_arrow_down
Yes. Every model we release is backed by a rigorous paper, and we continue submitting to top conferences like ICLR, EMNLP, SIGIR, NeurIPS, and ICML.
handshake
I'm not yet a Jina or Elastic customer, but I want to use the Reader API, model APIs, or cloud marketplace images. What should I do?
keyboard_arrow_down
Simply sign up and pay through our website or the relevant cloud marketplace, just as before.
handshake
Can I buy a commercial license for Jina models from Elastic?
keyboard_arrow_down
Yes. Since August 10, 2026, Elastic sells commercial licenses for Jina models directly, as their own SKU. This covers running Jina models in your own self-managed, on-premises, or air-gapped infrastructure, and it is available through Elastic direct, federal, and cloud service provider (CSP) channels. The offering is called Jina On-Prem. To get a quote, contact Elastic Sales.
handshake
What is Jina On-Prem?
keyboard_arrow_down
Jina On-Prem is a commercial license plus a set of self-contained Docker containers that let you run Jina models entirely inside your own infrastructure. The containers make no external connections: there is no call to Hugging Face or any model registry, and no license server, telemetry, or logging endpoint, which is what makes them viable in an air-gapped network. They cover the Jina model portfolio, including embedding, reranker, and reader models, and they expose Elastic Inference Service (EIS), OpenAI, Cohere, Voyage AI, and Gemini API schemas, so existing applications work without code changes. It is a separate SKU: it is not based on Elastic Resource Units (ERUs), and you do not need to run Elasticsearch to use it. It also works alongside open source Elasticsearch. It has been available to order since August 10, 2026; contact Elastic Sales for availability and terms.
handshake
How is Jina On-Prem priced?
keyboard_arrow_down
It is an annual license fee, scoped by which models you deploy and how many GPUs run inference for them. There is no per-token billing. Pricing is not self-serve: contact Elastic Sales for a quote for your deployment.
handshake
Who is Jina On-Prem for?
keyboard_arrow_down
Organizations that cannot, or prefer not to, send data to a cloud AI service. Typical cases are air-gapped and high-security environments, public sector and defense, regulated industries such as financial services and healthcare, latency-critical or offline systems, and teams that want a fixed, predictable inference cost instead of per-token pricing. If that describes your deployment, Elastic Sales can work through the details with you.
handshake
I'm an Elastic customer. Can I use Jina models in Elastic Cloud without deploying anything?
keyboard_arrow_down
Yes. Jina models are available through the Elastic Inference Service (EIS), so you can use them for ingest and search without provisioning machine learning nodes or managing GPU infrastructure. The models generally available on EIS include jina-embeddings-v5-text-small, jina-embeddings-v5-text-nano, jina-embeddings-v5-omni-small, jina-embeddings-v5-omni-nano, jina-embeddings-v3, jina-clip-v2, and the jina-reranker-v3.5, jina-reranker-v3, jina-reranker-v2-base-multilingual and jina-reranker-m0 rerankers. Consult the Elastic documentation for the current model list, supported regions, and the minimum stack version for each model.
handshake
I downloaded the weights from Hugging Face. Do I need a license to use them in production?
keyboard_arrow_down
If the model is licensed CC-BY-NC 4.0, yes. Non-commercial use, including evaluation and research, is free. Commercial production use is not covered by that license and requires a commercial license, which is what Jina On-Prem provides. Legacy Apache-2.0 models are not affected. If you are unsure which applies to your deployment, contact Elastic Sales or your Elastic account team.
handshake
I want to sign a contract or a custom agreement covering Jina models. What should I do?
keyboard_arrow_down
Contact Elastic Sales. Commercial licensing, contracting, and support for Jina models now run through Elastic's standard sales and support process.
handshake
I'm purchasing your services as a Chinese entity. Can I get a Chinese invoice (发票)?
keyboard_arrow_down
Self-serve API purchases are invoiced by Jina AI GmbH in Germany, so we cannot issue a Chinese invoice (发票) for them. For larger or contract-based purchases, contact Elastic Sales to discuss the available contracting entities.
handshake
I'm an Elastic customer and want to learn best practices for embeddings and rerankers, or I'm generally interested in Jina AI's development. What should I do?
keyboard_arrow_down
Contact Elastic Sales, and we can arrange a session between you, the Jina AI team, and Elastic.

How to get my API key?

video_not_supported

What's the rate limit?

Rate Limit
Rate limits are tracked in three ways: RPM (requests per minute), and TPM (tokens per minute). Limits are enforced per IP/API key and will be triggered when either the RPM or TPM threshold is reached first. When you provide an API key in the request header, we track rate limits by key rather than IP address.
ProductAPI EndpointDescriptionarrow_upwardw/o API Keykey_offw/ Free API Keykeyw/ Paid API Keykeyw/ Premium API KeykeyAverage LatencyToken Usage CountingAllowed Request
Reader APIhttps://r.jina.aiConvert URL to LLM-friendly text20 RPM500 RPM500 RPMtrending_up5000 RPM7.9sCount the number of tokens in the output response.GET/POST
Reader APIhttps://s.jina.aiSearch the web and convert results to LLM-friendly textblock100 RPM100 RPMtrending_up1000 RPM2.5sEvery request costs a fixed number of tokens, starting from 10000 tokensGET/POST
Embedding APIhttps://api.jina.ai/v1/embeddingsConvert text/images to fixed-length vectorsblock100 RPM & 100,000 TPM500 RPM & 2,000,000 TPMtrending_up5,000 RPM & 50,000,000 TPM
ssid_chart
depends on the input size
help
Count the number of tokens in the input request.POST
Reranker APIhttps://api.jina.ai/v1/rerankRank documents by queryblock100 RPM & 100,000 TPM500 RPM & 2,000,000 TPMtrending_up5,000 RPM & 50,000,000 TPM
ssid_chart
depends on the input size
help
Count the number of tokens in the input request.POST
Classifier APIhttps://api.jina.ai/v1/trainTrain a classifier using labeled examplesblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokens × num_itersPOST
Classifier API (Few-shot)https://api.jina.ai/v1/classifyClassify inputs using a trained few-shot classifierblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokensPOST
Classifier API (Zero-shot)https://api.jina.ai/v1/classifyClassify inputs using zero-shot classificationblock25 RPM & 25,000 TPM125 RPM & 500,000 TPM1,250 RPM & 12,000,000 TPM
ssid_chart
depends on the input size
Tokens counted as: input_tokens + label_tokensPOST
Segmenter APIhttps://api.jina.ai/v1/segmentTokenize and segment long text20 RPM200 RPM200 RPM1,000 RPM0.3sToken is not counted as usage.GET/POST
DeepSearchhttps://deepsearch.jina.ai/v1/chat/completionsReason, search and iterate to find the best answerblock50 RPM50 RPM500 RPM56.7sCount the total number of tokens in the whole process.POST

Do I need a commercial license?

CC BY-NC License Self-Check

play_arrow
Are you using our official API or official images on Azure, AWS, or GCP?
play_arrow
Yes
No restrictions. Simply sign up and pay through our website or cloud marketplace.
play_arrow
No
play_arrow
Are you a paid Elastic customer?
play_arrow
Yes
Commercial use is likely already included in your Elastic license. Contact your Elastic Sales representative if unsure.
Contact sales
play_arrow
No
We're currently unable to issue standalone commercial licensing agreements. Please contact Elastic Sales for more information.
Contact sales

Other questions

Reader-related common questions
What are the costs associated with using the Reader API?
keyboard_arrow_down
Reader is free for basic usage: prepend 'https://r.jina.ai/' to your URL. Supplying an API key raises the rate limit and charges tokens based on content length. See Q16 for rate limits.
How does the Reader API function?
keyboard_arrow_down
The Reader API fetches the URL server-side and returns clean, LLM-ready text. You choose the fetching engine with the X-Engine header: direct issues a plain HTTP fetch and is the fastest, the default engine renders the page in a headless browser so client-side JavaScript executes before extraction, and cf-browser-rendering is an experimental Cloudflare-backed renderer. Boilerplate such as navigation, headers, footers, and ads is stripped, and the main content is converted to Markdown. Use X-Respond-With to get other shapes of the same page, and X-Target-Selector or X-Remove-Selector to keep or drop specific CSS selectors.
Is the Reader API open source?
keyboard_arrow_down
The Reader service code is available on the Jina AI GitHub organization. The models used by Reader, including ReaderLM-v2 and jina-vlm, are licensed CC-BY-NC 4.0, which is not an open-source license: they are free to download and use non-commercially, but commercial production use requires a commercial license. Elastic has sold that license separately since August 10, 2026; contact Elastic Sales to arrange one.
What is the typical latency for the Reader API?
keyboard_arrow_down
It depends mainly on the engine and on the page itself. A direct fetch of a simple page typically returns in a few hundred milliseconds, while the default browser engine has to load and execute the page before extraction and usually lands in the low seconds. Heavy single-page apps, slow origin servers, and large PDFs take longer. Repeating the same URL within 5 minutes is served from cache and returns almost immediately, so a warm URL is much faster than a cold one.
Why should I use the Reader API instead of scraping the page myself?
keyboard_arrow_down
Scraping can be complicated and unreliable, particularly with complex or dynamic pages. The Reader API provides a streamlined, reliable output of clean, LLM-ready text.
Does the Reader API support multiple languages?
keyboard_arrow_down
The Reader API is language-agnostic and returns content in the original language of the page; it does not translate. Content negotiation is available through the X-Locale header, which sets the browser locale used when rendering, so sites that serve different markup per locale can be steered to the version you want.
Does the Reader API respect website access controls?
keyboard_arrow_down
Yes. Reader operates as a standard web client and respects website access controls. If a website blocks the request, that outcome is respected. You are responsible for ensuring your use of Reader complies with the terms of the sites you access and does not infringe third-party intellectual property rights.
Can the Reader API extract content from PDF files?
keyboard_arrow_down
Yes, the Reader API can natively extract content from PDF files.
Can the Reader API process media content from web pages?
keyboard_arrow_down
Yes, Reader can caption images on webpages using the x-with-generated-alt header. This adds descriptive alt tags to images that lack them, enabling LLMs to understand visual content. Video summarization is planned for future releases.
Is it possible to use the Reader API on local HTML files?
keyboard_arrow_down
No, the Reader API can only process content from publicly accessible URLs.
Does Reader API cache the content?
keyboard_arrow_down
If you request the same URL within 5 minutes, the Reader API will return the cached content.
Can I use the Reader API to access content behind a login?
keyboard_arrow_down
Yes, for pages that accept cookie-based sessions. Pass your session cookies with the X-Set-Cookie header and the Reader forwards them when fetching the URL, using the same <name>=<value> form as a normal Set-Cookie, optionally scoped with ; domain=. Requests carrying cookies are never cached, so each one is a fresh fetch. This does not perform a login for you: it replays credentials you already hold, so you are responsible for obtaining them and for complying with the target site's terms of service. Flows that require an interactive login, MFA, or a bearer token the page fetches itself are out of scope.
Can I use the Reader API to access PDF on arXiv?
keyboard_arrow_down
Yes, you can either use the native PDF support from the Reader (https://r.jina.ai/https://arxiv.org/pdf/2310.19923v4) or use the HTML version from the arXiv (https://r.jina.ai/https://arxiv.org/html/2310.19923v4)
How does image caption work in Reader?
keyboard_arrow_down
Reader captions all images at the specified URL and adds `Image [idx]: [caption]` as an alt tag (if they initially lack one). This enables downstream LLMs to interact with the images in reasoning, summarizing etc.
What is the scalability of the Reader? Can I use it in production?
keyboard_arrow_down
The Reader API is designed to be highly scalable. It is auto-scaled based on the real-time traffic and the maximum concurrency requests is now around 4000. We are maintaining it actively as one of the core products of Jina AI. So feel free to use it in production.
What is the rate limit of the Reader API?
keyboard_arrow_down
Please find the latest rate limit information in the table below. Note that we are actively working on improving the rate limit and performance of the Reader API, the table will be updated accordingly.
speedRate limit
What is ReaderLM? How can I use it?
keyboard_arrow_down
ReaderLM-v2 is a 1.54B parameter small language model that converts raw HTML into clean Markdown or JSON, and can extract structured data using a JSON schema or natural language instructions. Use it via the Reader API with the x-respond-with: readerlm-v2 header, or deploy it from the AWS, Azure, or GCP marketplaces. For image-heavy or scanned documents, jina-vlm is our 2.4B parameter vision-language reader model.
launchAWS SageMakerlaunchGoogle CloudlaunchMicrosoft Azure
How do I extract structured data from webpages?
keyboard_arrow_down
Use the x-json-schema header with a JSON schema definition, or use x-instruction header with natural language instructions. Both features work with ReaderLM-v2 to extract specific fields like prices, titles, dates, etc. from any webpage into structured JSON format.
Does Reader actively bypass website anti-bot protection?
keyboard_arrow_down
No. Reader does not actively circumvent or bypass any website defense mechanisms, anti-bot systems, or access controls. If a website detects our service as a bot and blocks the request, that outcome is respected. We operate as a standard web client and do not employ techniques designed to evade detection systems. You remain responsible for ensuring your use of Reader respects third-party intellectual property rights and the terms of the sites you access.
Will upgrading from a free to a paid API key give me access to more websites?
keyboard_arrow_down
No. Upgrading from a free tier to a paid API key does not grant access to additional websites or bypass any site restrictions. The difference between tiers is primarily in rate limits and performance optimizations. A paid API key provides higher request throughput and faster processing, but it does not enable access to websites that block our service.
Can I run Reader inside my own infrastructure?
keyboard_arrow_down
Yes. The reader models ship as self-contained offline Docker containers under the Jina On-Prem commercial license that Elastic has sold as its own SKU since August 10, 2026. This is the path for air-gapped and firewalled environments, where the containers make no outbound connections of any kind. Note that fetching arbitrary public web pages still requires network access to those pages; the on-prem value is in running the extraction models locally. Contact Elastic Sales for a quote.
Embeddings-related common questions
How were the Jina embedding models trained?
keyboard_arrow_down
For detailed information on our training processes, data sources, and evaluations, refer to the technical reports on arXiv. The jina-embeddings-v5 text models are trained in two stages: embedding distillation from a larger teacher model, followed by task-specific LoRA adapter training on frozen backbone weights. The v5-omni multimodal variants add a third stage that trains only cross-modal projectors, leaving the text backbone and adapters frozen.
launcharXiv
What are your multimodal embedding models?
keyboard_arrow_down
jina-embeddings-v5-omni-small (~1.74B parameters, 1024 dimensions, 32K context) and jina-embeddings-v5-omni-nano (~1.04B parameters, 768 dimensions, 8K context) are our current multimodal models. They accept text, images, audio, video, and PDFs in one shared vector space, so you can index in one modality and query in another without reindexing. Their text-only output is identical to jina-embeddings-v5-text-small and jina-embeddings-v5-text-nano respectively, which means you can add multimodal input to an existing text index without re-embedding it. jina-clip-v2 (865M parameters) remains available as a lighter text-and-image option.
launcharXiv
Which languages do your models support?
keyboard_arrow_down
All models released since 2024 are multilingual. jina-embeddings-v5-text-small and the v5-omni models are built on a Qwen3 backbone with broad multilingual coverage; jina-embeddings-v5-text-nano is built on EuroBERT-210M, covering 15 major European and global languages including English, French, German, Spanish, Chinese, Japanese, Arabic, and Hindi. jina-embeddings-v3 and jina-clip-v2 support 89 languages. For per-language benchmark numbers, see the MMTEB results in each model's technical report.
launcharXiv
What is the maximum context length for a single input?
keyboard_arrow_down
Context length varies by model: jina-embeddings-v5-text-small and jina-embeddings-v5-omni-small support up to 32,768 tokens, while jina-embeddings-v5-text-nano and jina-embeddings-v5-omni-nano support 8,192 tokens. jina-embeddings-v4 and the jina-code-embeddings models support 32,768 tokens; jina-embeddings-v3, jina-clip-v2, and jina-colbert-v2 support 8,192 tokens. Inputs above the limit return an error unless you set truncate: true.
What is the maximum number of inputs I can include in a single request?
keyboard_arrow_down
There is no hard limit on the number of items per request. The API batches inputs internally by token count for optimal GPU utilization, so you can send as many texts or images as needed in a single request. PDFs are the exception: send one PDF per request.
How do I send images, audio, video, or PDFs to the multimodal models?
keyboard_arrow_down
Pass a typed object in the input array using an image, audio, video, or pdf key, whose value is either a public URL or base64-encoded bytes. The model routes each modality to the appropriate encoder, and you can mix modalities freely within a single batch. Supported audio formats include WAV, MP3, FLAC, OGG, M4A, and Opus; video is processed as 32 uniformly sampled frames. PDFs must be sent one per request.
How do Jina embeddings compare to the latest OpenAI, Cohere, and Voyage models?
keyboard_arrow_down
jina-embeddings-v5-text-small (677M parameters) is the strongest model under 1B parameters on MMTEB, scoring 67.0 average at task level, and reaches 71.7 average on English MTEB. jina-embeddings-v5-text-nano (239M parameters) scores 65.5 on MMTEB, ahead of every model we evaluated under 500M parameters. Our design target is capability per parameter rather than raw size, so these models are cheaper to serve than most alternatives at comparable or better retrieval quality. All v5 models support Matryoshka Representation Learning, so you can truncate dimensions down to 32 without retraining.
How seamless is the transition from OpenAI's text-embedding-3-large to your solution?
keyboard_arrow_down
The transition is straightforward: our API endpoint matches the input and output JSON schemas of OpenAI's text-embedding-3-large, so in most codebases you change the base URL, the API key, and the model name. The same holds for the Jina On-Prem containers, which additionally expose Elastic Inference Service (EIS), Cohere, Voyage AI, and Gemini schemas, so Jina models are a drop-in replacement in existing code paths. Note that embeddings from different model families are not comparable, so you have to re-embed your corpus rather than mix vectors from two providers in one index. For the On-Prem containers themselves, contact Elastic Sales.
How are tokens calculated for images and other non-text inputs?
keyboard_arrow_down
Text is counted in the standard way. Non-text inputs are converted to tokens by the relevant encoder, and the cost depends heavily on which model you use, so measure with your own inputs rather than assuming. As a reference point, a 600x600 pixel image costs approximately: • jina-embeddings-v5-omni-small: ~363 tokens • jina-embeddings-v5-omni-nano: ~362 tokens • jina-embeddings-v4: ~4,840 tokens • jina-clip-v2: ~16,000 tokens The v5-omni models are one to two orders of magnitude cheaper per image than the older models. Every response includes a usage object with the exact token count for that request, including an image_tokens breakdown for multimodal inputs, so you can verify cost per call.
Do you provide models for embedding images, audio, or video?
keyboard_arrow_down
Yes. jina-embeddings-v5-omni-small and jina-embeddings-v5-omni-nano embed text, images, audio, video, and PDFs into a single shared vector space. jina-embeddings-v4 and jina-clip-v2 handle text and images.
Can Jina embedding models be fine-tuned on private or company data?
keyboard_arrow_down
There are two paths. The self-serve one is the Fine-tuning API, which generates synthetic training data from a description of your domain and returns a fine-tuned model, without your having to assemble a labelled dataset. For fine-tuning on proprietary data under a commercial agreement, on dedicated infrastructure, or on a model that is not in the base model selector, Contact Elastic Sales; that work is scoped and contracted through Elastic.
Contact
Can the models be hosted privately, on my own infrastructure or in my own cloud account?
keyboard_arrow_down
Yes, in two ways. Jina models are available on the AWS, Azure, and GCP marketplaces, so you can deploy them inside your own cloud account. For self-managed, on-premises, or air-gapped infrastructure, Elastic sells a commercial license called Jina On-Prem, available since August 10, 2026, which ships the models as fully offline Docker containers with no external calls and no license server. To get a quote for either path, contact Elastic Sales.
launchAWS SageMakerlaunchGoogle CloudlaunchMicrosoft Azure
What is the 'task' parameter and when should I use it?
keyboard_arrow_down
The task parameter selects a task-specific LoRA adapter for optimal performance. Use retrieval.query for search queries, retrieval.passage for documents being searched, text-matching for symmetric similarity such as duplicate or paraphrase detection, classification for categorization, and separation for clustering. Retrieval is asymmetric, so using the wrong side of the query/passage pair measurably degrades results. The parameter is supported by jina-embeddings-v5, jina-embeddings-v4, and jina-embeddings-v3.
What is late-interaction retrieval and which models support it?
keyboard_arrow_down
Late interaction keeps token-level vectors instead of collapsing a document into one vector, which preserves fine-grained detail at the cost of a larger index. jina-embeddings-v4 supports both dense (single-vector) and late-interaction (multi-vector) output via the output_type parameter, and jina-colbert-v2 is a dedicated late-interaction model. For most retrieval pipelines, a dense v5 model followed by a reranker is the better accuracy-per-cost tradeoff.
What is late chunking and when should I use it?
keyboard_arrow_down
Late chunking embeds the whole document first with a long-context model, then derives chunk embeddings from the token-level representations. Unlike naive chunking, which embeds each chunk in isolation, late chunking preserves cross-chunk context, which improves retrieval quality in RAG pipelines where a chunk refers to something defined earlier in the document. Enable it with the late_chunking parameter.
Why does the API enforce a different context length than the model supports?
keyboard_arrow_down
Some models are architecturally capable of longer context than the hosted API accepts. Very long sequences consume substantial GPU memory, and we tune the serving configuration to balance throughput, latency, and cost for the majority of use cases. If you need the full architectural context length, run the model yourself: contact Elastic Sales about a self-managed deployment.
Why is jina-embeddings-v4 free, and why is it slow?
keyboard_arrow_down
jina-embeddings-v4 is built on the Qwen2-VL base model, released under the Qwen Research License, which permits research and non-commercial use only. We therefore cannot license it commercially and provide it free of charge via the API instead. It is also a 3.8B parameter model, so it is inherently slower per request, and we throttle its throughput to manage infrastructure costs. It is not suitable for production workloads, and for the same reason it is not offered through the Elastic Inference Service or as part of Jina On-Prem. For production use, take the jina-embeddings-v5 family: it is faster, stronger on retrieval benchmarks, and can be licensed for commercial use through Elastic Sales.
What are the rate limits for the Embeddings API?
keyboard_arrow_down
Rate limits depend on your API key type:

Free: 100 RPM, 100K TPM
Paid: 500 RPM, 2M TPM
Premium: 5,000 RPM, 50M TPM

There is an additional IP-based limit of 10,000 requests per 60 seconds to prevent abuse. Limits are applied per key and counted over a 60-second window, so bursts are smoothed rather than queued; a request over the limit returns HTTP 429 and should be retried with exponential backoff. If you need limits beyond the Premium tier, or dedicated capacity with no shared-tenant limit at all, contact Elastic Sales.
Which embedding model should I choose?
keyboard_arrow_down
Start with jina-embeddings-v5-text-small for text retrieval: it is the strongest sub-1B model we ship and handles 32K context. Drop to jina-embeddings-v5-text-nano when latency, cost, or edge hardware matters more than the last point of accuracy. Use jina-embeddings-v5-omni-small or v5-omni-nano when images, audio, video, or PDFs are involved; their text output is identical to the corresponding text model, so you can add modalities to an existing index without re-embedding. Use jina-code-embeddings-0.5b or 1.5b for source code. Within a family, newer is better.
What are the file size limits for images and PDFs?
keyboard_arrow_down
Maximum file sizes are 5 MB for images and 8 MB for PDFs. Larger files are rejected with an error.
Reranker-related common questions
How much does the Reranker API cost?
keyboard_arrow_down
Reranker API pricing follows the same token-based structure as the Embeddings API, and tokens are shared across all Jina APIs on the same key. New API keys include free tokens to get started; beyond that, token packages are available for purchase. See the pricing section for details.
What are the differences between the Jina rerankers?
keyboard_arrow_down
jina-reranker-v3.5 is our current flagship: a 0.6B parameter multilingual listwise reranker with 131K context, and a drop-in replacement for jina-reranker-v3. It improves on v3 across every axis we measure, with the largest gains on structured-data and legal retrieval, and runs 1.22x to 1.56x faster. jina-reranker-v3 remains available. jina-reranker-m0 is the multimodal reranker for ranking visual documents. jina-reranker-v2-base-multilingual is a smaller cross-encoder supporting 100+ languages, useful when you need a non-Qwen-derived model. jina-colbert-v2 uses late interaction across 89 languages.
How are the Jina rerankers licensed?
keyboard_arrow_down
jina-reranker-v3.5, jina-reranker-v3, jina-reranker-m0, jina-reranker-v2-base-multilingual, and jina-colbert-v2 are released under CC-BY-NC 4.0. You are free to use, share, and adapt them for non-commercial purposes. Commercial production use requires a commercial license, which Elastic has sold as its own SKU since August 10, 2026 under the name Jina On-Prem. Contact Elastic Sales for a quote. Legacy jina-reranker-v1-* models remain Apache-2.0.
Do the rerankers support multiple languages?
keyboard_arrow_down
Yes, all current rerankers are multilingual. jina-reranker-v3.5 improves on v3 on MIRACL and on multilingual retrieval generally. jina-reranker-v3 and jina-reranker-v2-base-multilingual support 100+ languages, jina-reranker-m0 handles multilingual visual document ranking, and jina-colbert-v2 supports 89 languages.
What is the maximum context length for each reranker?
keyboard_arrow_down
Context length varies by model:

jina-reranker-v3.5: 131,072 tokens (query plus all documents combined) with auto-truncation
jina-reranker-v3: 131,072 tokens with auto-truncation
jina-reranker-m0: 10,000 tokens
jina-reranker-v2-base-multilingual: 1,024 tokens, with automatic chunking for longer documents
jina-colbert-v2: 8,192 tokens

For the v1 and v2 rerankers, queries are auto-truncated and long documents are chunked with max-pooling across chunks.
Is there a limit on the number of documents I can rerank per query?
keyboard_arrow_down
There is no hard limit on the number of documents per request. Like our Embeddings API, the Reranker API batches inputs internally by token count for optimal GPU utilization. You can send as many documents as needed in a single request.
What latency can I expect when reranking 100 documents?
keyboard_arrow_down
Latency varies from 100 milliseconds to 7 seconds, depending largely on the length of the documents and the query. For instance, reranking 100 documents of 256 tokens each with a 64-token query takes about 150 milliseconds. Increasing the document length to 4096 tokens raises the time to 3.5 seconds. If the query length is increased to 512 tokens, the time further increases to 7 seconds.
Below is the time cost of reranking one query and 100 documents in milliseconds:
Number of tokens in each document
Number of tokens in the query256512102420484096
64156323136621073571
128194369137721233598
256273475139721554299
5124681385211435367068
Can the rerankers be hosted privately, on my own infrastructure or in my own cloud account?
keyboard_arrow_down
Yes. The rerankers are available on the AWS, Azure, and GCP marketplaces for deployment in your own cloud account. For self-managed, on-premises, or air-gapped infrastructure, Elastic sells a commercial license (Jina On-Prem) that ships the models as fully offline Docker containers. Contact Elastic Sales for a quote.
launchAWS SageMakerlaunchGoogle CloudlaunchMicrosoft Azure
Do you offer a reranker fine-tuned on domain-specific data?
keyboard_arrow_down
Before commissioning a fine-tune, try jina-reranker-v3.5: it was trained with self-distillation specifically for domain robustness and shows large gains on legal and structured-data retrieval over v3. A well-chosen off-the-shelf reranker plus better chunking usually closes more of the gap than a fine-tune does, and it costs nothing to test. If it still falls short on your data, a domain-specific reranker is a custom engagement: contact Elastic Sales to scope it.
Contact
What's the minimum image size for the documents?
keyboard_arrow_down
The minimum acceptable image size for the jina-reranker-m0 model is 28x28 pixels.
What is listwise reranking and how does it differ from pointwise?
keyboard_arrow_down
jina-reranker-v3 and jina-reranker-v3.5 use a listwise architecture: the query and all candidates share one context window and are scored in a single forward pass, so the model can compare documents against each other. Traditional pointwise rerankers, including jina-reranker-v2-base-multilingual, score each document independently against the query. Listwise scoring is more accurate because relevance is often relative to what else is in the candidate set.
Why does the API enforce a different context length than the model supports?
keyboard_arrow_down
Some rerankers are architecturally capable of longer context than the hosted API accepts. Very long sequences consume substantial GPU memory, and we tune the serving configuration to balance throughput, latency, and cost for the majority of use cases. If you need the full architectural context length, run the model in your own infrastructure and contact Elastic Sales about a commercial license.
What are the rate limits for the Reranker API?
keyboard_arrow_down
Rate limits depend on your API key type:

Free: 100 RPM, 100K TPM
Paid: 500 RPM, 2M TPM
Premium: 5,000 RPM, 50M TPM

There is also an IP-based limit of 10,000 requests per 60 seconds. The same limits apply to the Embeddings and Reranker APIs, and tokens are shared across all Jina APIs on the same key.
Which reranker should I choose?
keyboard_arrow_down
Use jina-reranker-v3.5 for text. It is a drop-in replacement for jina-reranker-v3: the request schema is unchanged, so switching the model string is the entire migration. Use jina-reranker-m0 when your candidates are images or visually rich documents. Use jina-reranker-v2-base-multilingual when you need a smaller model or one that is not derived from a Qwen backbone.
API-related common questions
code
Can I use the same API key across all Jina APIs?
keyboard_arrow_down
Yes. One API key is valid for all Jina AI search foundation products, including the Reader, Embeddings, Reranker, Classifier, and Segmenter APIs, with tokens shared across all of them.
code
Can I monitor the token usage of my API key?
keyboard_arrow_down
Yes, token usage can be monitored in the 'API Key & Billing' tab by entering your API key, allowing you to view the recent usage history and remaining tokens. If you have logged in to the API dashboard, these details can also be viewed in the 'Manage API Key' tab.
code
What should I do if I forget my API key?
keyboard_arrow_down
If you have misplaced a topped-up key and wish to retrieve it, please contact support AT jina.ai with your registered email for assistance. It's recommended to log in to keep your API key securely stored and easily accessible.
Contact
code
Do API keys expire?
keyboard_arrow_down
No, our API keys do not have an expiration date. If a key is compromised, revoke it yourself in the API Key Management dashboard, which takes effect immediately; issue a replacement key first if you want to avoid downtime. Any remaining token balance stays on your account rather than on the revoked key. If you cannot access the dashboard, or believe the account itself is compromised, raise it with Elastic Support.
Contact
code
Can I transfer tokens between API keys?
keyboard_arrow_down
Yes, you can transfer tokens from a premium key to another. After logging into your account on the API Key Management dashboard, use the settings of the key you want to transfer out to move all remaining paid tokens.
code
Can I revoke my API key?
keyboard_arrow_down
Yes, you can revoke your API key if you believe it has been compromised. Revoking a key will immediately disable it for all users who have stored it, and all remaining balance and associated properties will be permanently unusable. If the key is a premium key, you have the option to transfer the remaining paid balance to another key before revocation. Notice that this action cannot be undone. To revoke a key, go to the key settings in the API Key Management dashboard.
code
Why is the first request for some models slow?
keyboard_arrow_down
This is because our serverless architecture offloads certain models during periods of low usage. The initial request activates or 'warms up' the model, which may take a few seconds. After this initial activation, subsequent requests process much more quickly.
code
Is my API data used to train your models?
keyboard_arrow_down
No. We never use your API requests, inputs, or outputs to train our embedding, reranker, or any other models. Your data remains yours.
code
What are the rate limits for Jina APIs?
keyboard_arrow_down
Rate limits apply per API key:

Free: 100 RPM, 100K TPM
Paid: 500 RPM, 2M TPM
Premium: 5,000 RPM, 50M TPM

There is also an IP-based limit of 10,000 requests per 60 seconds. Limits vary by endpoint; see the rate limit table above for per-endpoint figures.
code
Are there batch size limits for the APIs?
keyboard_arrow_down
There is no batch size limit for either the Embeddings or Reranker APIs. You can send as many items or documents as needed per request. Both APIs batch inputs internally by token count for optimal GPU utilization.
code
Are the Jina APIs the same thing as Jina models inside Elastic?
keyboard_arrow_down
No, they are three separate paths. The Jina APIs on this site are self-serve and pay-as-you-go with a Jina API key. The Elastic Inference Service (EIS) runs Jina models inside Elastic Cloud, billed through your Elastic subscription, with no infrastructure for you to manage. Jina On-Prem is a commercial license, sold by Elastic as its own SKU since August 10, 2026, for running the models in your own self-managed, on-premises, or air-gapped infrastructure. For the EIS and On-Prem paths, contact Elastic Sales.
Billing-related common questions
attach_money
Is billing based on the number of sentences or requests?
keyboard_arrow_down
Our pricing model is based on the total number of tokens processed, allowing users the flexibility to allocate these tokens across any number of sentences, offering a cost-effective solution for diverse text analysis requirements.
attach_money
Is there a free trial available for new users?
keyboard_arrow_down
Yes. New users get an auto-generated API key with free tokens usable across any of our models. Once the free tokens are consumed, you can purchase additional tokens for the key in the 'Buy tokens' tab.
attach_money
Are tokens charged for failed requests?
keyboard_arrow_down
No, tokens are not deducted for failed requests.
attach_money
What payment methods are accepted?
keyboard_arrow_down
Payments are processed through Stripe, supporting a variety of payment methods including credit cards, Google Pay, and PayPal for your convenience.
attach_money
Is invoicing available for token purchases?
keyboard_arrow_down
For self-serve token purchases, Stripe issues an invoice to the email address associated with your Stripe account at the time of purchase. If you need a formal purchase order, a negotiated contract, procurement paperwork, or consolidated billing, that runs through Elastic rather than Stripe: contact Elastic Sales.
attach_money
How do I buy a commercial license rather than API tokens?
keyboard_arrow_down
Token purchases on this site cover use of the hosted Jina APIs. They do not license you to run the model weights in your own infrastructure. For that, Elastic has sold a commercial license as its own SKU since August 10, 2026, priced annually rather than per token. Contact Elastic Sales for a quote.
attach_money
Can I pay by invoice or purchase order instead of card?
keyboard_arrow_down
Self-serve token purchases are processed through Stripe and invoiced automatically to your Stripe account email. For purchase orders, procurement processes, or volumes above what self-serve top-up supports, contact Elastic Sales.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.