Overview
BGE-reranker-v2-m3 is a lightweight cross-encoder built on BGE-M3. Unlike an embedding model it takes the query and the passage together and outputs a relevance score directly, which is more accurate than comparing independently-encoded vectors. BAAI positions it as the multilingual, fast-inference, easy-to-deploy member of the v2 reranker family. Raw scores are unbounded logits; a sigmoid maps them to 0-1.
Strengths
- Multilingual reranking at roughly 568M parameters
- Fast enough for the online path
- Same lineage as bge-m3, the default embedder
Use cases
- Second-stage reranking after vector search
- Raising RAG precision cheaply
Good to know
Every candidate costs a forward pass, so rerank a shortlist rather than a corpus. Outputs are logits, not probabilities.
API access
One OpenAI-compatible API for every model in the catalog, plus the CLI and the function SDK inside workflows. Sign up free, add a card for $5 in credits, and these requests work:
# Reranker models are called through the # platform APIs and workflows; list what is deployed: curl https://gateway.graphn.ai/v1/models \ -H "Authorization: Bearer $GRAPHN_API_KEY"
graphn model get bge-reranker
@function
async def search(kb_id: str, query: str) -> list:
return await kb.search(kb_id, query, top_k=5, rerank=True)