Elastic
Jina AI
模型
API
keyboard_arrow_down
Reader
把任意 URL 轉成 Markdown,為大模型提供更好的事實依據。
向量模型
多模態多語言向量模型。
重排模型
讓搜尋相關性最大化的重排模型。
MCP
terminal
命令列
article
llms.txt
smart_toy
智慧體
data_object
Schema
menu_book
文件
登入
login
新聞稿
二月 28, 2024

運用多任務對比學習革新雙語文本嵌入

我們的新論文探討了我們的西班牙語-英語和德語-英語模型如何運用多任務對比學習和複雜的數據管線,來掌握長達 8192 個 tokens 的文本的語言理解和跨語言效率
Composite image of four colorful, stylized landmarks: Brandenburg Gate, St. Peter's Basilica, Tiananmen, and Golden Gate Brid
Jina AI • 3 分鐘閱讀

在我們最近發表的論文 Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings 中,我們詳細介紹了 德語-英語和西班牙語-英語雙語文本嵌入模型的開發過程。

Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
We introduce a novel suite of state-of-the-art bilingual text embedding models that are designed to support English and another target language. These models are capable of processing lengthy text inputs with up to 8192 tokens, making them highly versatile for a range of natural language processing tasks such as text retrieval, clustering, and semantic textual similarity (STS) calculations. By focusing on bilingual models and introducing a unique multi-task learning objective, we have significantly improved the model performance on STS tasks, which outperforms the capabilities of existing multilingual models in both target language understanding and cross-lingual evaluation tasks. Moreover, our bilingual models are more efficient, requiring fewer parameters and less memory due to their smaller vocabulary needs. Furthermore, we have expanded the Massive Text Embedding Benchmark (MTEB) to include benchmarks for German and Spanish embedding models. This integration aims to stimulate further research and advancement in text embedding technologies for these languages.
arXiv.orgIsabelle Mohr
Embedding API
Start with 1M free tokens. Top-performing, 8192 context length bilingual embeddings for your search and RAG systems.

我們的方法採用多任務對比學習和先進的數據處理流程,專注於雙語能力的同時,將token長度擴展至 8192。這種方法使我們的模型在理解目標語言和進行跨語言評估時都能高效運作。

Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.
GitHub
Ich bin ein Berliner: German-English Bilingual Embeddings with 8K Token Length
Jina AI introduces a German/English bilingual embedding model, featuring an extensive 8,192-token length, specifically designed to support German businesses thriving in the U.S. market.
GitHub

除了論文中提到的雙語模型外,我們還開發了中文-英語雙語模型和英語單語模型。這些擴展展示了我們致力於滿足廣泛語言需求並進一步提升語言處理能力的承諾。

8K Token-Length Bilingual Embeddings Break Language Barriers in Chinese and English
The first bilingual Chinese-English embedding model with 8192 token-length.
Discord
Jina AI Launches World's First Open-Source 8K Text Embedding, Rivaling OpenAI
Jina AI introduces jina-embeddings-v2, the world's first open-source model boasting an 8K context length. Matching the prowess of OpenAI's proprietary models, this innovation is now publicly accessible on Huggingface, signaling a significant milestone in the landscape of text embeddings.

我們的雙語模型以其高效性為特色,通過優化詞彙量來減少參數數量和記憶體使用。這種效率突顯了我們致力於創建既強大又資源高效的語言處理工具的決心。

在論文發布後,我們擴展了 Massive Text Embedding Benchmark (MTEB),加入了英語-德語和英語-西班牙語嵌入模型的基準測試。這項擴展是我們努力推動非英語語言文本嵌入技術研究和發展的一部分。

在 Jina AI,我們的目標是提升多語言的處理和理解能力,通過開發雙語和單語文本嵌入模型為自然語言處理領域做出貢獻。

類別:
新聞稿
rss_feed

閱讀更多
八月 03, 2026 • 11 分鐘閱讀
jina-reranker-v3.5: Faster Listwise Reranking with Hybrid Attention and Self-Distillation
Jina AI
五月 12, 2026 • 7 分鐘閱讀
jina-embeddings-v5-omni: Embeddings for Text, Image, Audio and Video
Han Xiao
二月 19, 2026 • 7 分鐘閱讀
jina-embeddings-v5-text: New SOTA Small Multilingual Embeddings
Han Xiao
Abstract digital artwork in black and white, featuring scattered dots forming letters in a halftone effect. The central lette
當前語言 / 主題
搜尋底座
Reader
向量模型
重排模型
獲取 Jina API 金鑰
速率限制
關於我們
新聞
下載 Jina 標誌
open_in_new
下載 Elastic 標誌
open_in_new
API 狀態
Elastic © 2026.安全條款及條件隱私管理 Cookie請勿出售或分享我的個人資訊
本網站及其所有相關內容、軟體、產品和服務僅供專業使用,不面向消費者。