Qdrant Summer of Code 2024 - ONNX Cross Encoders in Python
Huong (Celine) Hoang
·October 14, 2024

On this page:
Editor’s note: The score tests check numerical agreement between model implementations. This 2024 post was edited for length and clarity. Read the original version.
I’m Huong (Celine) Hoang, and I worked on cross-encoder reranking in FastEmbed during Qdrant Summer of Code 2024. FastEmbed already generated embeddings for retrieval. My project added models that score a query together with each candidate document, so applications could reorder retrieved results.
This post records my 2024 internship. For current models and runnable examples, use the FastEmbed reranker documentation.
Adding a Different Model Output
Embedding models turn text into vectors. Cross-encoders instead take query-document pairs and return relevance scores. That required a new input-output scheme in FastEmbed.
I worked on TextCrossEncoderBase and OnnxCrossEncoder, drawing on the existing embedding classes. The interface needed to hide model loading and tokenization while giving users access to the scores.
The project included models such as Xenova/ms-marco-MiniLM-L-6-v2, Xenova/ms-marco-MiniLM-L-12-v2, and BAAI rerankers. We used Open Neural Network Exchange (ONNX) models so FastEmbed could run them without requiring PyTorch or TensorFlow.
FastEmbed scores the candidates retrieved by Qdrant; the application orders them by those scores.
Tokenization and Model Integration
A cross-encoder processes the query and document together. I configured paired-input tokenization and, for models that use them, token type IDs to distinguish the two inputs. Model-specific configurations needed care because the supported models did not all tokenize inputs in the same way.
Loading an ONNX model was only part of the work. The runtime inputs, batching, and output handling also had to agree with the original model. I added batching support so users could score multiple candidates through the same interface.
Checking the Scores
We compared ONNX outputs with the corresponding PyTorch models to check the conversions. Tests used the same query and documents in each implementation, then compared the resulting scores.
One test used the query “What is the capital of France?” with documents about Paris and Berlin. It checked FastEmbed’s scores against saved PyTorch outputs with an absolute tolerance of 1e-3.
Model configurations, tokenizers, and tests all needed debugging. My mentor, George Panchuk, helped me work through those issues and keep the code readable during review.
Lessons From the Internship
The project added cross-encoder support for the FastEmbed 0.4.0 release. At the end of the internship, possible follow-ups included more models, improved batch processing, and additional tokenizer support.
The main lesson for me was to test the complete user experience. A model that loads successfully still needs correct inputs, useful outputs, and an interface developers can understand.
Thank you to George and the Qdrant team for their guidance. To try reranking in an application, follow the current FastEmbed reranker example.