在我們最近發表的論文 Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings 中,我們詳細介紹了 德語-英語和西班牙語-英語雙語文本嵌入模型的開發過程。
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
We introduce a novel suite of state-of-the-art bilingual text embedding models that are designed to support English and another target language. These models are capable of processing lengthy text inputs with up to 8192 tokens, making them highly versatile for a range of natural language processing tasks such as text retrieval, clustering, and semantic textual similarity (STS) calculations. By focusing on bilingual models and introducing a unique multi-task learning objective, we have significantly improved the model performance on STS tasks, which outperforms the capabilities of existing multilingual models in both target language understanding and cross-lingual evaluation tasks. Moreover, our bilingual models are more efficient, requiring fewer parameters and less memory due to their smaller vocabulary needs. Furthermore, we have expanded the Massive Text Embedding Benchmark (MTEB) to include benchmarks for German and Spanish embedding models. This integration aims to stimulate further research and advancement in text embedding technologies for these languages.

Embedding API
Start with 1M free tokens. Top-performing, 8192 context length bilingual embeddings for your search and RAG systems.

我們的方法採用多任務對比學習和先進的數據處理流程,專注於雙語能力的同時,將token長度擴展至 8192。這種方法使我們的模型在理解目標語言和進行跨語言評估時都能高效運作。
Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.

Ich bin ein Berliner: German-English Bilingual Embeddings with 8K Token Length
Jina AI introduces a German/English bilingual embedding model, featuring an extensive 8,192-token length, specifically designed to support German businesses thriving in the U.S. market.

除了論文中提到的雙語模型外,我們還開發了中文-英語雙語模型和英語單語模型。這些擴展展示了我們致力於滿足廣泛語言需求並進一步提升語言處理能力的承諾。
8K Token-Length Bilingual Embeddings Break Language Barriers in Chinese and English
The first bilingual Chinese-English embedding model with 8192 token-length.

Jina AI Launches World's First Open-Source 8K Text Embedding, Rivaling OpenAI
Jina AI introduces jina-embeddings-v2, the world's first open-source model boasting an 8K context length. Matching the prowess of OpenAI's proprietary models, this innovation is now publicly accessible on Huggingface, signaling a significant milestone in the landscape of text embeddings.

我們的雙語模型以其高效性為特色,通過優化詞彙量來減少參數數量和記憶體使用。這種效率突顯了我們致力於創建既強大又資源高效的語言處理工具的決心。
在論文發布後,我們擴展了 Massive Text Embedding Benchmark (MTEB),加入了英語-德語和英語-西班牙語嵌入模型的基準測試。這項擴展是我們努力推動非英語語言文本嵌入技術研究和發展的一部分。
在 Jina AI,我們的目標是提升多語言的處理和理解能力,通過開發雙語和單語文本嵌入模型為自然語言處理領域做出貢獻。







