新聞稿
二月 28, 2024

運用多任務對比學習革新雙語文本嵌入

我們的新論文探討了我們的西班牙語-英語和德語-英語模型如何運用多任務對比學習和複雜的數據管線,來掌握長達 8192 個 tokens 的文本的語言理解和跨語言效率
Jina AI • 3 分鐘閱讀

在我們最近發表的論文 Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings 中,我們詳細介紹了 德語-英語和西班牙語-英語雙語文本嵌入模型的開發過程。

Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
We introduce a novel suite of state-of-the-art bilingual text embedding models that are designed to support English and another target language. These models are capable of processing lengthy text inputs with up to 8192 tokens, making them highly versatile for a range of natural language processing tasks such as text retrieval, clustering, and semantic textual similarity (STS) calculations. By focusing on bilingual models and introducing a unique multi-task learning objective, we have significantly improved the model performance on STS tasks, which outperforms the capabilities of existing multilingual models in both target language understanding and cross-lingual evaluation tasks. Moreover, our bilingual models are more efficient, requiring fewer parameters and less memory due to their smaller vocabulary needs. Furthermore, we have expanded the Massive Text Embedding Benchmark (MTEB) to include benchmarks for German and Spanish embedding models. This integration aims to stimulate further research and advancement in text embedding technologies for these languages.
Embedding API
Start with 1M free tokens. Top-performing, 8192 context length bilingual embeddings for your search and RAG systems.

我們的方法採用多任務對比學習和先進的數據處理流程,專注於雙語能力的同時,將token長度擴展至 8192。這種方法使我們的模型在理解目標語言和進行跨語言評估時都能高效運作。

Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.
Ich bin ein Berliner: German-English Bilingual Embeddings with 8K Token Length
Jina AI introduces a German/English bilingual embedding model, featuring an extensive 8,192-token length, specifically designed to support German businesses thriving in the U.S. market.

除了論文中提到的雙語模型外,我們還開發了中文-英語雙語模型和英語單語模型。這些擴展展示了我們致力於滿足廣泛語言需求並進一步提升語言處理能力的承諾。

8K Token-Length Bilingual Embeddings Break Language Barriers in Chinese and English
The first bilingual Chinese-English embedding model with 8192 token-length.
Jina AI Launches World's First Open-Source 8K Text Embedding, Rivaling OpenAI
Jina AI introduces jina-embeddings-v2, the world's first open-source model boasting an 8K context length. Matching the prowess of OpenAI's proprietary models, this innovation is now publicly accessible on Huggingface, signaling a significant milestone in the landscape of text embeddings.

我們的雙語模型以其高效性為特色,通過優化詞彙量來減少參數數量和記憶體使用。這種效率突顯了我們致力於創建既強大又資源高效的語言處理工具的決心。

在論文發布後,我們擴展了 Massive Text Embedding Benchmark (MTEB),加入了英語-德語和英語-西班牙語嵌入模型的基準測試。這項擴展是我們努力推動非英語語言文本嵌入技術研究和發展的一部分。

在 Jina AI,我們的目標是提升多語言的處理和理解能力,通過開發雙語和單語文本嵌入模型為自然語言處理領域做出貢獻。

類別:
新聞稿

閱讀更多
九月 14, 2026 • 9 分鐘閱讀
jina-ocr-v1:在低預算 GPU 上實現更快速的文件解析
八月 03, 2026 • 11 分鐘閱讀
jina-reranker-v3.5:透過混合注意力機制與自我知識蒸餾實現更快速的列表式重排器
五月 12, 2026 • 7 分鐘閱讀
jina-embeddings-v5-omni:支援文字、圖片、音訊與影片的向量模型