← Managed inference

BGE Reranker v2

Cheap multilingual reranker for a retrieval shortlist

Overview

BGE-reranker-v2-m3 is a lightweight cross-encoder built on BGE-M3. Unlike an embedding model it takes the query and the passage together and outputs a relevance score directly, which is more accurate than comparing independently-encoded vectors. BAAI positions it as the multilingual, fast-inference, easy-to-deploy member of the v2 reranker family. Raw scores are unbounded logits; a sigmoid maps them to 0-1.

Strengths

  • Multilingual reranking at roughly 568M parameters
  • Fast enough for the online path
  • Same lineage as bge-m3, the default embedder

Use cases

  • Second-stage reranking after vector search
  • Raising RAG precision cheaply

Good to know

Every candidate costs a forward pass, so rerank a shortlist rather than a corpus. Outputs are logits, not probabilities.

API access

One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:

bge-reranker · curl
# Reranker models are called through the
# platform APIs and workflows; list what is deployed:
curl https://gateway.graphn.ai/v1/models \
  -H "Authorization: Bearer $GRAPHN_API_KEY"