Elastic
Jina AI
Models
API
keyboard_arrow_down
Reader
Convert any URL to Markdown for better grounding LLMs.
Embeddings
Multimodal multilingual embeddings.
Reranker
Reranker for maximizing search relevance.
MCP
terminal
CLI
article
llms.txt
smart_toy
Agents
data_object
Schema
menu_book
Docs
Log in
login
Reader
copyright CC BY-NC 4.0
open_in_new Release Post

ReaderLM-v2

Frontier small language model for converting raw HTML into markdown or JSON
License
copyright CC-BY-NC-4.0
Release Date
calendar_month
2025-01-16
Input
abc
Text (HTML)
arrow_forward
Output
abc
Text (Markdown)
abc
Text (JSON)
Model Details
Parameters: 1.54B
Input Token Length: 512K
Base Model help_outline
open_in_new
Qwen2.5-1.5B-Instruct
Trained Languages help_outline
14 languages
Supported Languages help_outline
29 languages
Related Models
link
reader-lm-1.5b
Available via
Jina API
AWS SageMaker
Microsoft Azure
Google Cloud
Hugging Face
Air-gapped
I/O graph 1

HTML

ReaderLM-v2

Markdown

I/O graph 2

HTML

ReaderLM-v2

Instruction

JSON

I/O graph 3

HTML

ReaderLM-v2

Instruction

Markdown

Choose models to compare
Publications (1)
ICLR 2025
March 04, 2025
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON

Overview

ReaderLM-v2 is a 1.54B-parameter small language model that transforms messy HTML into clean Markdown or JSON with high accuracy, processing documents up to 512K tokens. It is the ideal tool for grounding large language models: converting web content into structured formats that LLMs can consume efficiently. The model outperforms Qwen2.5-32B-Instruct and Gemini2-flash-expr on HTML-to-Markdown tasks while running at a fraction of their cost.

Methods

The model's effectiveness results from two key innovations. First, a three-stage data synthesis pipeline generates high-quality, diverse training data by iteratively drafting, refining, and critiquing web content extraction — this synthetic data approach ensures the model sees a wide variety of HTML structures without requiring manual annotation. Second, a unified training framework combines continuous pre-training with multi-objective optimization, allowing the model to learn both HTML-to-Markdown and HTML-to-JSON conversion simultaneously. The 'shallow-but-wide' decoder-only architecture (28 layers, 1536 hidden dimensions, 12 query heads, 2 KV heads) is optimized for selective-copy operations. The 512K token context is enabled through zigzag-ring-attention. Contrastive loss training significantly reduces degeneration issues.

Performance

On HTML-to-Markdown tasks, the model achieves ROUGE-L of 0.84, Jaro-Winkler of 0.82, and Levenshtein distance of 0.22 — outperforming Qwen2.5-32B-Instruct and Gemini2-flash-expr. On HTML-to-JSON tasks, it maintains competitive performance with F1 scores of 0.81 and 98% pass rate. The model processes at 67 tokens/s input and 36 tokens/s output on a T4 GPU. Degeneration issues (token loops, repetitive output) are significantly reduced through contrastive loss training. The 512K token context window eliminates the need for chunking on most real-world documents.

Best Practice

The model is accessible through a Google Colab notebook demonstrating HTML-to-Markdown conversion, JSON extraction, and instruction-following. For HTML-to-Markdown tasks, input raw HTML without prefix instructions. For JSON extraction, specify the target schema in the prompt. The create_prompt helper function facilitates easy prompt creation for both tasks. The model works on Colab's free T4 GPU tier (requires vllm and triton) but has limitations without bfloat16 or Flash Attention 2 support; RTX 3090/4090 is recommended for production. Available on AWS SageMaker, Azure, and GCP marketplace. Licensed under CC BY-NC 4.0 for non-commercial use. Use this model as a preprocessing step in RAG pipelines to convert web content into clean Markdown before embedding with jina-embeddings-v5-text-small.

Blogs that mention this model
May 25, 2025 • 21 minutes read
What We Learned at ICLR2025
We collect some most interesting papers in ICLR 2025, featuring TIPS, FlexPrefill, Zero-Shot Rerankers, SVD-LLM, Hymba etc.
Jina AI
Three people smiling on a stage at a conference with an ICLR banner visible, suggesting a warm and lively event atmosphere.
May 07, 2025 • 9 minutes read
Model Soup’s Recipe for Embeddings
Boost robustness and performance with model soups: averaging weights. No extra cost, better results.
Bo Wang
Scott Martens
Still life drawing of a purple bowl filled with apples and oranges on a white table. The scene features rich colors against a
April 08, 2025 • 21 minutes read
jina-reranker-m0: Multilingual Multimodal Document Reranker
Introducing jina-reranker-m0, our new multilingual multimodal reranker for retrieving visual documents, with SOTA performance on multilingual long documents and code searching tasks.
Jina AI
Modern dot matrix text display on a dark blue background, conveying a digital feel.
January 31, 2025 • 14 minutes read
A Practical Guide to Deploying Search Foundation Models in Production
We offer detailed cost and performance breakdowns for three deployment strategies: Jina API, self-hosted K8s, and AWS SageMaker, to help you make the right decision.
Saahil Ognawala
Scott Martens
Abstract cityscape illustration with orange, grey and white buildings, featuring visible balconies with a potted plant.
January 15, 2025 • 17 minutes read
ReaderLM v2: Frontier Small Language Model for HTML to Markdown and JSON
ReaderLM-v2 is a 1.5B small language model for HTML-to-Markdown conversion and HTML-to-JSON extraction with exceptional quality.
Jina AI
Orange text "ReaderLM-u2" on a vibrant dark red digital screen.
Current language / theme
Search Foundation
Reader
Embeddings
Reranker
Get Jina API key
Rate Limit
About us
News
Download Jina logo
open_in_new
Download Elastic logo
open_in_new
API Status
Elastic © 2026.SecurityTerms & ConditionsPrivacyManage CookiesDo Not Sell or Share My Personal Information
This website and all associated content, software, products, and services are intended for professional use only. No consumer use is intended or directed.