Auto Fine-Tuning
Just tell us which domain you want your embeddings to excel in, and we automatically deliver a ready-to-use, fine-tuned embedding model for that domain.
What is Auto Fine-Tuning?
There are three ways to specify your requirement: a general instruction, a URL, or a query-document description. Choose one.
public
Or, webpage URL
Refer to the content from a URL for fine-tuning.
keyboard_arrow_down
notes
Or, general instruction
Provide a detailed description of how the fine-tuned embeddings will be used.
keyboard_arrow_down
Select a base embedding model
Fine-tuning allows you to take a pre-trained model and adapt it to a specific task or domain by training it on a new dataset. In practice, finding effective training data is not straightforward for many users. Effective training requires more than just throwing raw PDFs, HTMLs into the model; and it is hard to get it right. Auto fine-tuning solves this problem by automatically generating effective training data using an advanced LLM agent pipeline; and fine-tuning the model within a ML workflow. You can think it as a combination of synthetic data generation and AutoML, so all you need to do is describe your target domain in natural language and let our system do the rest.
Auto fine-tuning holds an auto-magical promise to deliver fine-tuned embeddings for any domain you want. But does it really work? This is a fairly reasonable doubt. We've tested it on a variety of domains and base models to find out. Check out the cherry-picked and lemon-picked results below.
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-enAvg. improvement
arrow_upward 2%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.505 arrow_forward 0.532 arrow_upward 5%
MAP
0.352 arrow_forward 0.389 arrow_upward 10%
MRR
0.352 arrow_forward 0.389 arrow_upward 10%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from tollefj/norwegian-nli-triplets
NDCG
0.852 arrow_forward 0.867 arrow_upward 2%
MAP
0.800 arrow_forward 0.820 arrow_upward 2%
MRR
0.800 arrow_forward 0.820 arrow_upward 2%
data_usage Synthetic data generated
Total
4648
Training
4480
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-enAvg. improvement
arrow_upward 6%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.672 arrow_forward 0.755 arrow_upward 12%
MAP
0.567 arrow_forward 0.675 arrow_upward 19%
MRR
0.567 arrow_forward 0.675 arrow_upward 19%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from mteb/askubuntudupquestions-reranking
NDCG
0.698 arrow_forward 0.722 arrow_upward 3%
MAP
0.515 arrow_forward 0.549 arrow_upward 6%
MRR
0.666 arrow_forward 0.712 arrow_upward 7%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-enAvg. improvement
arrow_upward 9%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.727 arrow_forward 0.861 arrow_upward 18%
MAP
0.640 arrow_forward 0.814 arrow_upward 27%
MRR
0.640 arrow_forward 0.814 arrow_upward 27%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from mteb/scidocs-reranking
NDCG
0.773 arrow_forward 0.822 arrow_upward 6%
MAP
0.575 arrow_forward 0.651 arrow_upward 13%
MRR
0.823 arrow_forward 0.884 arrow_upward 7%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-zhAvg. improvement
arrow_upward 1%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.718 arrow_forward 0.785 arrow_upward 9%
MAP
0.629 arrow_forward 0.717 arrow_upward 14%
MRR
0.629 arrow_forward 0.717 arrow_upward 14%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from C-MTEB/CMedQAv2-reranking
NDCG
0.938 arrow_forward 0.948 arrow_upward 1%
MAP
0.912 arrow_forward 0.926 arrow_upward 2%
MRR
0.920 arrow_forward 0.933 arrow_upward 1%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-enAvg. improvement
arrow_upward 6%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.543 arrow_forward 0.579 arrow_upward 7%
MAP
0.402 arrow_forward 0.452 arrow_upward 12%
MRR
0.402 arrow_forward 0.452 arrow_upward 12%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from nc33/triplet_sbert_law2 (machine-translated to dutch)
NDCG
0.904 arrow_forward 0.948 arrow_upward 5%
MAP
0.870 arrow_forward 0.930 arrow_upward 7%
MRR
0.870 arrow_forward 0.930 arrow_upward 7%
data_usage Synthetic data generated
Total
9128
Training
8960
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-codeAvg. improvement
arrow_downward -4%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.671 arrow_forward 0.640 arrow_downward -5%
MAP
0.569 arrow_forward 0.525 arrow_downward -8%
MRR
0.569 arrow_forward 0.525 arrow_downward -8%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from mteb/stackoverflowdupquestions-reranking
NDCG
0.640 arrow_forward 0.621 arrow_downward -3%
MAP
0.530 arrow_forward 0.505 arrow_downward -5%
MRR
0.555 arrow_forward 0.532 arrow_downward -4%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-codeAvg. improvement
arrow_downward -4%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.632 arrow_forward 0.711 arrow_upward 13%
MAP
0.517 arrow_forward 0.622 arrow_upward 20%
MRR
0.517 arrow_forward 0.622 arrow_upward 20%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from mteb/stackoverflowdupquestions-reranking
NDCG
0.640 arrow_forward 0.619 arrow_downward -3%
MAP
0.530 arrow_forward 0.504 arrow_downward -5%
MRR
0.555 arrow_forward 0.525 arrow_downward -5%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Base model for fine-tuning
jinaai/jina-embeddings-v2-base-enAvg. improvement
arrow_upward 1%
description Domain instruction
speed Performance on synthetic validation set before and after fine-tuning
NDCG
0.646 arrow_forward 0.729 arrow_upward 13%
MAP
0.535 arrow_forward 0.644 arrow_upward 20%
MRR
0.535 arrow_forward 0.644 arrow_upward 20%
speed Performance on held-out test set before and after fine-tuning
check
Tested on 50 random samples from mteb/askubuntudupquestions-reranking
NDCG
0.645 arrow_forward 0.650 arrow_upward 1%
MAP
0.452 arrow_forward 0.462 arrow_upward 2%
MRR
0.606 arrow_forward 0.605 arrow_downward -0%
data_usage Synthetic data generated
Total
616
Training
448
Validation
168
Auto Fine-Tuning API
Get fine-tuned embeddings for any domain you want.
globe_book
Use
r.jina.ai to read a URL and fetch its contenttravel_explore
Use
s.jina.ai to search the web and get SERPAdd
mcp.jina.ai as your MCP server to use our APIs in LLMsContent Format
Return-Format
You can control the level of detail in the response to prevent over-filtering. The default pipeline is optimized for most websites and LLM input.
Default
arrow_drop_down
JSON Response
Accept
The response will be in JSON format, containing the URL, title, content, and timestamp (if available). In Search mode, it returns a list of five entries, each following the described JSON structure.
Timeout (seconds)
X-Timeout
Maximum time to wait for page load. Increase for slow pages, decrease for simple static pages.
Token Budget
X-Token-Budget
Limits the maximum number of tokens used for this request. Exceeding this limit will cause the request to fail.
Use jina-ocr-v1
X-Respond-With
Uses jina-ocr-v1 to convert images to Markdown, delivering high-quality results on pages and documents with complex structure and content. Costs 40× more tokens!open_in_newLearn more
Pagination
X-Page
The page to process. Focuses on the specified page of a multi-page document, and only applies when a document such as a PDF is being processed.
Extract Only (CSS Selector)
X-Target-Selector
Only extract content matching these CSS selectors. Example: article, .main-content, #post-body
Wait For (CSS Selector)
X-Wait-For-Selector
Wait until these elements appear before extracting content. Useful for dynamically loaded content.
Exclude (CSS Selector)
X-Remove-Selector
Remove these elements before extraction. Example: nav, footer, .sidebar, #ads
Remove All Images
X-Retain-Images
Strip all images from the output. Reduces token usage when images are not needed.
OpenAI Citation Format
X-Retain-Links
Format links for OpenAI's web browsing tool. Uses special citation markers compatible with GPT models.open_in_newLearn more
Links Summary Section
X-With-Links-Summary
A "Buttons & Links" section is added at the end. This helps downstream LLMs and web agents navigate the page or take further actions.
None
arrow_drop_down
Images Summary Section
X-With-Images-Summary
An "Images" section will be created at the end. This gives the downstream LLMs an overview of all visuals on the page, which may improve reasoning.
None
arrow_drop_down
Browser Viewport Size
POST
viewport
Set browser window dimensions. Affects responsive layouts and content visibility.open_in_newLearn more
Forward Cookie
X-Set-Cookie
Our API server can forward your custom cookie settings when accessing the URL, which is useful for pages requiring extra authentication. Note that requests with cookies will not be cached.open_in_newLearn more
Image Caption
X-With-Generated-Alt
Captions all images at the specified URL, adding 'Image [idx]: [caption]' as an alt tag for those without one. This allows downstream LLMs to interact with the images in activities such as reasoning and summarizing.
Use a Proxy Server
X-Proxy-Url
Our API server can utilize your proxy to access URLs, which is helpful for pages accessible only through specific proxies.open_in_newLearn more
Use a Country-Specific Proxy Server
X-Proxy
Set country code for location-based proxy server. Use 'auto' for optimal selection or 'none' to disable.
Bypass Cached Content
X-No-Cache
Our API caches URL contents for a certain amount of time. Set it to true to ignore the cached result and fetch the content from the URL directly.
Cache Tolerance (seconds)
X-Cache-Tolerance
Accept cached content if younger than N seconds. Set to 0 for fresh content (same as Bypass Cache), or higher values to allow faster responses from cache.
Page Ready Timing
X-Respond-Timing
When to consider a page fully loaded. Later timings wait longer but capture more dynamic content.
Default
arrow_drop_down
Custom User-Agent
X-User-Agent
Override the browser User-Agent string. Useful for accessing sites that require specific browsers or block crawlers.
Custom Referer
X-Referer
Set the HTTP Referer header. Some sites check this to verify traffic comes from expected sources.
Preserve Base64 Images
X-Keep-Img-Data-Url
Keep inline base64-encoded images in markdown output instead of converting them to external URLs.
Do Not Cache or Track
DNT
Prevent this request from being cached or logged on our servers. Use for sensitive URLs.
GitHub Flavored Markdown
X-No-Gfm
Opt in/out features from GFM (GitHub Flavored Markdown).
Enabled
arrow_drop_down
Stream Mode
Accept
Stream mode is beneficial for large target pages, allowing more time for the page to fully render. If standard mode results in incomplete content, consider using Stream mode.open_in_newLearn more
Customize Browser Locale
X-Locale
Control the browser locale to render the page. Lots of websites serve different content based on the locale.open_in_newLearn more
Respect robots.txt
X-Robots-Txt
Check robots.txt rules before fetching. Specify which bot name to use for the check.
Include iframe Content
X-With-Iframe
Extract content from embedded iframes. Enable for pages with content loaded in iframes.
Include Shadow DOM
X-With-Shadow-Dom
Extract content from Shadow DOM components. Enable for pages using web components.
Use Final URL as Base
X-Base
Resolve relative URLs using the final destination URL after redirects, instead of the original URL.
Local PDF/HTML file
POST
Use Reader on local PDF and HTML files by uploading them. Only PDF and HTML files are supported. For HTML, also specify a reference URL so related CSS/JS scripts can be parsed correctly.
upload
Run JavaScript Before Extraction
POST
Execute custom JS to modify the page before content extraction. Can be inline code or a URL to a script file.open_in_newLearn more
Heading Style
X-Md-Heading-Style
Sets markdown heading format (passed to Turndown).
Hash Style
arrow_drop_down
Horizontal Rule Style
X-Md-Hr
Defines markdown horizontal rule format (passed to Turndown).
Bullet Point Style
X-Md-Bullet-List-Marker
Sets bullet list marker character (passed to Turndown).
*
arrow_drop_down
Emphasis Style
X-Md-Em-Delimiter
Defines markdown emphasis delimiter (passed to Turndown).
_
arrow_drop_down
Strong Emphasis Style
X-Md-Strong-Delimiter
Sets markdown strong emphasis delimiter (passed to Turndown).
**
arrow_drop_down
Link Style
X-Md-Link-Style
Determines markdown link format (passed to Turndown).
Inline
arrow_drop_down
EU Residency
Experimental
When enabled, infrastructure and data processing for this request are located within the EU.
curl "https://r.jina.ai/https://www.example.com"key
API key
visibility_off
Available tokens
0
code
Can I use the same API key across all Jina APIs?
keyboard_arrow_down
code
Can I monitor the token usage of my API key?
keyboard_arrow_down
code
What should I do if I forget my API key?
keyboard_arrow_down
code
Do API keys expire?
keyboard_arrow_down
code
Can I transfer tokens between API keys?
keyboard_arrow_down
code
Can I revoke my API key?
keyboard_arrow_down
code
Why is the first request for some models slow?
keyboard_arrow_down
code
Is my API data used to train your models?
keyboard_arrow_down
code
What are the rate limits for Jina APIs?
keyboard_arrow_down
code
Are there batch size limits for the APIs?
keyboard_arrow_down
code
Are the Jina APIs the same thing as Jina models inside Elastic?
keyboard_arrow_down
Rate limit
Rate limits are tracked in two ways: RPM (requests per minute) and TPM (tokens per minute). Limits are enforced per IP/API key and will be triggered when either the RPM or TPM threshold is reached first. When you provide an API key in the request header, we track rate limits by key rather than IP address.
Columns
arrow_drop_down
| Product | API Endpoint | Descriptionarrow_upward | w/o API Keykey_off | w/ Free API Keykey | w/ Paid API Keykey | w/ Premium API Keykey | Average latency | Token Usage Counting | Allowed Request | |
|---|---|---|---|---|---|---|---|---|---|---|
| Reader API | https://r.jina.ai | Converts a URL to LLM-friendly text | 20 RPM | 500 RPM | 500 RPM | trending_up5000 RPM | 7.9s | Count the number of tokens in the output response. | GET/POST | |
| Reader API | https://s.jina.ai | Search the web and convert results to LLM-friendly text | block | 100 RPM | 100 RPM | trending_up1000 RPM | 2.5s | Every request costs a fixed number of tokens, starting from 10,000 tokens | GET/POST | |
| Embedding API | https://api.jina.ai/v1/embeddings | Convert text/images to fixed-length vectors | block | 100 RPM & 100,000 TPM | 500 RPM & 2,000,000 TPM | trending_up5,000 RPM & 50,000,000 TPM | ssid_chart depends on the input size help | Count the number of tokens in the input request. | POST | |
| Reranker API | https://api.jina.ai/v1/rerank | Rank documents by query | block | 100 RPM & 100,000 TPM | 500 RPM & 2,000,000 TPM | trending_up5,000 RPM & 50,000,000 TPM | ssid_chart depends on the input size help | Count the number of tokens in the input request. | POST |
Billing-related common questions
attach_money
Is billing based on the number of sentences or requests?
keyboard_arrow_down
attach_money
Is there a free trial available for new users?
keyboard_arrow_down
attach_money
Are tokens charged for failed requests?
keyboard_arrow_down
attach_money
What payment methods are accepted?
keyboard_arrow_down
attach_money
Is invoicing available for token purchases?
keyboard_arrow_down
attach_money
How do I buy a commercial license rather than API tokens?
keyboard_arrow_down
attach_money
Can I pay by invoice or purchase order instead of card?
keyboard_arrow_down
attach_money
I paid, but my balance or rate limit has not changed. What should I check?
keyboard_arrow_down
attach_money
How do I cancel, stop auto top-up, or remove a saved payment method?
keyboard_arrow_down