Best open source embedding models for search and RAG

7 minUpdated:
Best open source embedding models for search and RAG

Strong open source embedding families include BGE (including multilingual BGE-M3), E5, GTE, Nomic Embed and small sentence-transformers models such as all-MiniLM. Choose by language coverage, context length, vector size and licence, then confirm the choice on your own queries rather than leaderboard rank alone.

What does an embedding model do in RAG?

An embedding model turns text into a vector so that similar meanings land close together. In RAG, you embed document chunks once, store them in a vector database, then embed each query and fetch the nearest chunks.

Retrieval quality caps answer quality. If the right chunk is not retrieved, even the best generator cannot answer correctly, so the embedding model deserves as much attention as the LLM.

Embeddings are also useful outside RAG: semantic search boxes, deduplication, clustering support tickets and recommending related articles all use the same vectors.

That makes the embedding choice a platform decision. Once several features depend on one index, changing the model means coordinated re-indexing, so it pays to choose carefully the first time.

Which open source embedding models should you consider?

Model versions change often, so treat this as a map of families rather than a fixed ranking. Always check the model card for current size, context length and licence.

FamilyMaintainerLanguagesStrengthsBest forTrade-off
BGE / BGE-M3BAAIEnglish, Chinese; M3 is multilingualDense retrieval; M3 also offers sparse and multi-vector modesGeneral and multilingual RAGLarger variants need a GPU for fast indexing
E5MicrosoftEnglish and multilingual variantsSolid retrieval with simple query and passage prefixesSearch over mixed contentForgetting the prefixes hurts results
GTEAlibabaEnglish and multilingual variantsGood accuracy for its sizeBalanced quality and speedCheck the licence per variant
Nomic EmbedNomicMainly English, plus multilingual versionsLong context, open training data and codeLong chunks and transparencyNeeds task prefixes
all-MiniLM and similarsentence-transformersMainly EnglishVery small and fast on CPUPrototypes, edge and low-cost searchLower accuracy on hard queries
Jina embeddingsJina AIEnglish and multilingualLong context optionsLong documentsSome versions use non-commercial licences; check the licence file

How do you pick the right embedding model?

  • Language: if your users or documents are not English, start with a multilingual model such as BGE-M3 or multilingual E5.
  • Context length: match the model’s maximum input to your chunk size so text is not silently truncated.
  • Vector size: larger dimensions cost more storage and memory in the index; some models support shortened vectors.
  • Hardware: small models index on CPU; larger ones want a GPU for bulk embedding.
  • Licence: confirm commercial use is allowed for the exact checkpoint you download.
  • Evaluation: build 50 to 200 real queries with known relevant chunks and measure recall at k.

Are leaderboards like MTEB enough?

MTEB is the standard public benchmark for embeddings and a good way to shortlist. It is not a guarantee. Scores average many tasks, and your domain, language and query style may look nothing like the benchmark data.

Some models are also tuned in ways that favour benchmark tasks. A model a few places lower can win on your support tickets or legal clauses. Your own recall numbers are the tiebreaker.

Look at the task categories behind a score, not only the average. For RAG, the retrieval columns matter far more than clustering or classification results.

How do you serve embedding models?

The simplest path is the sentence-transformers Python library inside your ingestion job. For a shared service, Hugging Face Text Embeddings Inference provides an optimised HTTP server, and Ollama can serve several embedding models locally.

Batch document embedding offline, and keep query embedding on a low-latency path. Queries are short, so even a CPU can often handle them for moderate traffic.

Record the model name and version next to every stored vector. When you upgrade later, that metadata tells you exactly which documents still need re-embedding.

Watch the tokenizer. Every model counts length in its own tokens, so a chunk that fits one model may be truncated by another.

Where embedding-based retrieval breaks

  • Exact identifiers such as SKUs, error codes and names, which dense vectors blur. Add keyword search (BM25) and combine results.
  • Mixing vectors from two different models in the same index; they are not comparable.
  • Changing the model without re-embedding the whole corpus.
  • Skipping the query or passage prefixes that models like E5 and Nomic expect.
  • Chunks too long for the model’s context, silently cut off at the end.
  • No reranker, so the top result is merely close rather than truly relevant.

Should you add a reranker?

Usually yes. A cross-encoder reranker, such as the BGE reranker family, scores each query and candidate pair together and reorders the top results. It is slower than vector search, so you apply it to a short list, often the top 20 to 50.

The pattern of hybrid retrieval plus reranking is often worth more than swapping to a bigger embedding model. RepoLoot’s catalog flags RAG repositories that already implement hybrid search and reranking, which saves wiring it yourself.

How should you evaluate embeddings on your own data?

Build a small retrieval test set before choosing. Take real questions from users or support logs and mark which chunks answer each one. Fifty well-labelled queries tell you more than any public leaderboard.

Then embed your corpus with each candidate model, run the queries and measure recall at 5 and 10: how often the correct chunk appears in the top results. Also record index size and embedding time, because they drive running cost.

  • Include short keyword queries and long natural-language questions
  • Include queries in every language your users write in
  • Add a few questions with no answer to test false matches
  • Re-run the test whenever you change chunking

How do vector size and quantisation affect cost?

Every stored vector costs memory in your vector database, and memory is usually the largest line in the bill. A model with larger vectors multiplies that cost across millions of chunks.

Some models are trained so that you can truncate vectors to fewer dimensions with modest quality loss. Many vector databases also support scalar or binary quantisation. Test both on your recall set; the savings can be large while the quality drop is often small.

Which vector database pairs well with open embeddings?

Any of them. Qdrant, Weaviate, Milvus and pgvector all store vectors from any model, as long as the dimension matches the index. Choose the database for operations and filtering needs, and the model for retrieval quality, as two separate decisions.

If you already run Postgres and your corpus is modest, pgvector keeps everything in one database. Dedicated vector databases earn their place at larger scale or with heavy filtering and hybrid search needs.

Which embedding model should you start with?

Start with one of these, build your evaluation set, and only then try alternatives. Without the test set, switching models is guesswork.

Revisit the choice once or twice a year. New releases arrive frequently, and a quick run of your recall test tells you whether an upgrade is worth a re-index.

  • English-only prototype on a laptop: a small sentence-transformers model such as all-MiniLM.
  • Production English search with a GPU for indexing: a mid-sized BGE, E5 or GTE model.
  • Multilingual content or users: BGE-M3 or a multilingual E5 variant.
  • Long documents with few chunks: a long-context model such as Nomic Embed, after checking its prefixes.
  • Strict licensing requirements: shortlist only models with permissive licences on the exact checkpoint.

Frequently asked questions

What is the best open source embedding model?
It depends on language, domain and hardware. BGE-M3 is a strong multilingual default, E5 and GTE are reliable general choices, and small sentence-transformers models are best when you need CPU speed. Shortlist from MTEB, then measure recall on your own queries.
Are open source embeddings as good as paid APIs?
For many retrieval tasks, leading open models are competitive with hosted embedding APIs, and they avoid per-token fees and data leaving your servers. Results vary by domain, so compare both on a labelled sample of your real queries before deciding.
Can I change embedding models later?
Yes, but you must re-embed every document with the new model, because vectors from different models are not comparable. Keep raw chunks stored separately from vectors, and design ingestion so a full re-index is a routine job rather than a migration project.
Do I need a GPU for embeddings?
Not always. Small models run acceptably on CPU for query embedding and modest corpora. A GPU helps a lot when embedding large document collections or when using bigger models. Many teams embed in bulk on a rented GPU and serve queries on CPU.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides