보도 자료
2월 28, 2024

멀티태스크 대조 학습으로 혁신하는 이중 언어 텍스트 임베딩

새로운 논문에서는 스페인어-영어 및 독일어-영어 모델이 멀티태스크 대조 학습과 정교한 데이터 파이프라인을 활용하여 최대 8192 토큰의 텍스트에 대한 언어 이해와 교차 언어 효율성을 어떻게 달성하는지 살펴봅니다
Jina AI • 3 분 소요

최근 논문 Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings에서 우리는 독일어-영어와 스페인어-영어 이중 언어 텍스트 임베딩 모델의 개발 과정을 상세히 설명했습니다.

Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
We introduce a novel suite of state-of-the-art bilingual text embedding models that are designed to support English and another target language. These models are capable of processing lengthy text inputs with up to 8192 tokens, making them highly versatile for a range of natural language processing tasks such as text retrieval, clustering, and semantic textual similarity (STS) calculations. By focusing on bilingual models and introducing a unique multi-task learning objective, we have significantly improved the model performance on STS tasks, which outperforms the capabilities of existing multilingual models in both target language understanding and cross-lingual evaluation tasks. Moreover, our bilingual models are more efficient, requiring fewer parameters and less memory due to their smaller vocabulary needs. Furthermore, we have expanded the Massive Text Embedding Benchmark (MTEB) to include benchmarks for German and Spanish embedding models. This integration aims to stimulate further research and advancement in text embedding technologies for these languages.
Embedding API
Start with 1M free tokens. Top-performing, 8192 context length bilingual embeddings for your search and RAG systems.

우리의 접근 방식은 다중 작업 대조 학습과 고급 데이터 큐레이션 파이프라인을 활용하여 이중 언어 기능에 중점을 두면서 8192 토큰 길이까지 지원하도록 확장했습니다. 이 방법을 통해 우리의 모델은 대상 언어를 이해하고 교차 언어 평가를 효율적으로 수행하는 데 탁월한 성능을 보여줍니다.

Aquí Se Habla Español: Top-Quality Spanish-English Embeddings and 8k Context
Jina AI's new bilingual Spanish-English embedding model brings the state-of-the-art in AI to half a billion Spanish speakers.
Ich bin ein Berliner: German-English Bilingual Embeddings with 8K Token Length
Jina AI introduces a German/English bilingual embedding model, featuring an extensive 8,192-token length, specifically designed to support German businesses thriving in the U.S. market.

논문에서 다룬 이중 언어 모델 외에도 이중 언어 중국어-영어와 단일 언어 영어 모델도 개발했습니다. 이러한 추가 개발은 광범위한 언어적 요구를 충족시키고 언어 처리 능력을 향상시키려는 우리의 노력을 보여줍니다.

8K Token-Length Bilingual Embeddings Break Language Barriers in Chinese and English
The first bilingual Chinese-English embedding model with 8192 token-length.
Jina AI Launches World's First Open-Source 8K Text Embedding, Rivaling OpenAI
Jina AI introduces jina-embeddings-v2, the world's first open-source model boasting an 8K context length. Matching the prowess of OpenAI's proprietary models, this innovation is now publicly accessible on Huggingface, signaling a significant milestone in the landscape of text embeddings.

우리의 이중 언어 모델은 최적화된 어휘 크기로 작동하여 더 적은 매개변수와 메모리를 필요로 하는 효율성이 특징입니다. 이러한 효율성은 강력하면서도 자원 효율적인 언어 처리 도구를 만들려는 우리의 노력을 보여줍니다.

논문 발표 이후, 우리는 Massive Text Embedding Benchmark (MTEB)를 확장하여 영어-독일어와 영어-스페인어 임베딩 모델에 대한 벤치마크를 포함시켰습니다. 이러한 확장은 비영어권 언어에 대한 텍스트 임베딩 기술의 추가 연구와 발전을 촉진하기 위한 우리의 노력의 일환입니다.

Jina AI에서는 이중 언어 및 단일 언어 텍스트 임베딩 모델 개발을 통해 NLP 분야에 기여하면서 다중 언어의 처리와 이해를 향상시키는 것을 목표로 하고 있습니다.

범주:
보도 자료

자세히 보기
9월 14, 2026 • 9 분 소요
jina-ocr-v1: Faster Document Parsing on Low-Budget GPUs
8월 03, 2026 • 11 분 소요
jina-reranker-v3.5: Faster Listwise Reranking with Hybrid Attention and Self-Distillation
5월 12, 2026 • 7 분 소요
jina-embeddings-v5-omni: Embeddings for Text, Image, Audio and Video