Decide whether you need a reranker and how to evaluate it.
—
Plugsky is OpenAI-compatible with flat-rate plans — see pricing or start on the free plan (2 free models, no card).
The Reranker Comparison helps you evaluate reranking models and APIs for search and RAG. It focuses on the metrics that change retrieval quality: ranking accuracy on your queries, the latency added per search, and cost at your volume, so you can decide whether a cross-encoder pass is worth adding. It is for developers tuning hybrid search or RAG pipelines. Plugsky serves the generation side through an OpenAI-compatible endpoint.
A reranker scores each query and passage pair with a cross-encoder and reorders the candidate list, so the most relevant passages reach the model first. It is slower per item than vector search but far more accurate on the top results.
Usually yes for RAG, where a better top-k reduces wrong answers and prompt tokens. Measure the added milliseconds against the drop in irrelevant context and decide per use case; latency-sensitive search may skip it.
Retrieve a wide first-stage set, often 20 to 50 passages, rerank them, then keep the best 3 to 8 for the prompt. Going wider raises recall but costs more compute, so tune the pool size against your quality target.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs