> Explore Qdrant's agent skills catalog at https://skills.qdrant.tech/
> Search the documentation at https://skills.qdrant.tech/search?query=your+query+here
> Use this file to discover all available pages: https://qdrant.tech/llms.txt
We rewrote cross-encoder inference in Rust, benchmarked it against Python, and on a small reranking model it came out only 1.1x faster at the median. That is not a sign of a slow Rust implementation. The Python libraries we compared against, `fastembed` and `sentence-transformers`, hand the same work to the same C++ engine, `onnxruntime`, so all three spend almost all of their time in the same compiled code.

In the Rust community, rewriting code in Rust is called "oxidizing" it, since rust is also what iron turns into when it oxidizes. This post covers how we oxidized cross-encoder inference into a library, [`cross-encode-rs`](https://github.com/AstraBert/cross-encode-rs), how it works, where it does pull ahead (about 1.4x on a larger model), and where it does not (model load time).

## Understanding the Foundations

Before we dive into the implementation, two concepts will come up frequently:

- [**Cross-encoder**](https://www.sbert.net/examples/cross_encoder/applications/README.html): A model that reads a query and a document together and returns one relevance score for the pair. Cross-encoders are typically used for reranking: a fast first stage retrieves a pool of candidates, and the cross-encoder re-sorts them so the best matches come first. An embedding model (a bi-encoder) turns the query and each document into separate vectors and compares them afterward. A cross-encoder feeds both texts into the model at once, so every word of the query is compared with every word of the document. That makes it more accurate, and also more expensive: nothing can be computed ahead of time, so every query-document pair costs a full model run.
- [**ONNX Runtime**](https://onnxruntime.ai/docs/): ONNX (Open Neural Network Exchange) is a file format for trained models. You export a model from PyTorch or TensorFlow once, and run it anywhere. ONNX Runtime is the C++ engine that runs these files, with graph optimizations, hardware acceleration, and bindings for many languages, including Rust. We used the [`ort`](https://ort.pyke.io/) crate to load and run models.

With the fundamentals in place, let's look at how cross-encoder inference works and how we implemented it in Rust.

## How the Inference Works

Inference starts with two texts: a query and a document. These get fused into a single input built from three parallel arrays: token IDs, token types (query or document), and an attention mask marking which tokens are real and which are padding, so the model ignores the padding.

Once the input is ready, it runs through a few layers:

1. **Embedding**: Each token gets three embeddings, one for the token itself, one for its position, and one for its type (query or document). They are summed and normalized into a `[seq_len, hidden_dim]` matrix.
2. **Encoder stack**: Each layer runs multi-head self-attention, where every token attends to every other token in the pair. This is where the query and the document get compared. A small feed-forward network follows in each layer.
3. **Pooling**: The model keeps a single vector, usually the one at the `[CLS]` position (some models average all positions instead). Because of attention, this vector already summarizes the whole pair.
4. **Classification**: A linear layer maps that vector to `num_labels` outputs. Most cross-encoders use `num_labels = 1` and apply a sigmoid to get a score between 0 and 1. With `num_labels = 2`, softmax gives the probability of the "relevant" class.

The `ort` crate abstracts away all four layers. Our job in Rust is to tokenize the query and documents, batch them, build the token IDs, token types, and attention mask, run inference with `ort`, and apply sigmoid or softmax depending on `num_labels`.

![A query and a document are tokenized into one sequence with input IDs, token type IDs, and an attention mask. ort runs embedding, a six-layer encoder, pooling, and a classifier that returns one logit, and cross-encode-rs applies sigmoid to get the relevance score.](/blog/oxidizing-cross-encoders/cross-encoder-inference.svg)

*For the query `rust programming`, MiniLM scores `cargo builds rust code` at 0.8841 and `iron rusts in wet air` at 0.0035, although both documents contain the same `rust` token.*


Here is a brief overview in pseudo-Rust:

```rust
// tokenize text
 let encodings = tokenizer.encode_batch(
     vec![
         (query, document1),
         (query, document2)
      ]
);

// turn encodings into token IDs, attention mask
// and token types
let mask = encodings.get_attention_mask();
let ids = encodings.get_ids();
let types = encoding.get_type_ids();

// Convert our flattened arrays
// into 2-dimensional tensors
// of shape [num_documents, padding_dim]
let a_ids = TensorRef::from_array_view(
    ([num_documents, padding_dim], &*ids)
)?;
// ... same with mask and types ...

// run inference
let outputs = ort_model.run(
    ort::inputs![a_ids, a_mask, a_type_ids]
)?;

// extract logits into a 2D array
let logits = outputs[0]
    .try_extract_array::<f32>()?
    .into_dimensionality::<Ix2>()
    .unwrap();

// get num_labels
(_, num_labels) = logits.dim();

if num_labels == 1 {
    // apply sigmoid
} else {
    // apply softmax
}
```

<aside role="status">
The original code can be found <a href="https://github.com/AstraBert/cross-encode-rs/blob/main/crates/cross-encode-rs/src/inference.rs">here</a>.
</aside>

That's all it takes to oxidize cross-encoder inference. The main optimization in our Rust library is batching, which happens on two levels:

- All documents are packed together and encoded in one pass, rather than sequentially one-by-one.
- Input documents are sorted by tokenized length and run in batches of similar-sized documents, so computation isn't wasted on excess padding.

We also use the newly released `tokenizers` v1, which gives a modest speedup over v0.x.

## Comparing with Python

### Inference Latency

Two libraries dominate ONNX inference in the Python ecosystem: [`sentence-transformers`](https://sbert.net) and [`fastembed`](https://github.com/qdrant/fastembed). We compared our Rust crate, `cross-encode-rs`, against both using the [`mteb/scidocs-reranking`](https://huggingface.co/datasets/mteb/scidocs-reranking) dataset (`test` split, about 3.98K queries).

The benchmark design was simple: for each query in the dataset, we combined its positive and negative documents into one array and ran inference on them (30 query-document pairs per query, on average).

We ran the benchmark against two models and measured both per-request and per-document latency:

- [`Xenova/ms-marco-MiniLM-L-6-v2`](https://huggingface.co/Xenova/ms-marco-MiniLM-L-6-v2), a small cross-encoder built to run in the browser.
- [`jinaai/jina-reranker-v2-base-multilingual`](https://huggingface.co/jinaai/jina-reranker-v2-base-multilingual), a larger multilingual reranker.

We benchmarked all three libraries on both.

All benchmarks ran on a MacBook M4 Max with 48 GB of RAM, using all 14 CPU cores.

<aside role="status">
Note: <code>sentence-transformers</code> uses <code>torch</code> as its backend by default. For a fair comparison against the other two libraries, we configured it to use ONNX (via the <code>optimum</code> library) instead, pinned to <code>CPUExecutionProvider</code>. On macOS, Optimum otherwise picks CoreML first: in our first run, that put its MiniLM p50 at 213 ms per request, against 29 ms on the CPU provider.
</aside>

The results do not fit the usual "Python is slow" narrative:

- **On the small MiniLM model, the three libraries are close.** `cross-encode-rs` is about 1.1x faster at the median and 1.4x faster at p99, while the Python libraries are slightly faster on the quickest requests.
- **On the larger Jina model, `cross-encode-rs` pulls ahead.** It is 1.3 to 1.5x faster than both Python libraries at every percentile except the single slowest request.
- **The two Python libraries perform almost identically**, within 5% of each other from minimum to p99 on both models.

See the charts below for exact numbers.

All three libraries run on `onnxruntime`, the ONNX engine that Microsoft distributes across languages, which explains why their performance is so close on the small model. On the larger model, the two Python libraries still match each other while `cross-encode-rs` pulls about 1.4x ahead, so the runtime is not the whole story: how each library configures and feeds it also counts.

**Xenova/ms-marco-MiniLM-L-6-v2: inference latency**

| statistic | library | per_request_ms | per_document_ms |
| --- | --- | --- | --- |
| min | cross-encode-rs | 16.63 | 0.55 |
| min | fastembed | 13.79 | 0.46 |
| min | sentence-transformers | 13.71 | 0.46 |
| p50 | cross-encode-rs | 25.43 | 0.85 |
| p50 | fastembed | 27.45 | 0.92 |
| p50 | sentence-transformers | 28.73 | 0.96 |
| p90 | cross-encode-rs | 30.47 | 1.02 |
| p90 | fastembed | 36.16 | 1.21 |
| p90 | sentence-transformers | 37.53 | 1.26 |
| p99 | cross-encode-rs | 35.91 | 1.20 |
| p99 | fastembed | 49.23 | 1.64 |
| p99 | sentence-transformers | 51.43 | 1.72 |
| max | cross-encode-rs | 53.89 | 1.80 |
| max | fastembed | 71.80 | 2.39 |
| max | sentence-transformers | 71.25 | 2.37 |

_Per request: On MiniLM, the three libraries stay within 1.5x of each other at every percentile: at p99, a request takes 36 ms with cross-encode-rs, 49 ms with fastembed, and 51 ms with sentence-transformers._

_Per document: At p99, cross-encode-rs takes 1.20 ms per document, against 1.64 ms for fastembed and 1.72 ms for sentence-transformers._


**jinaai/jina-reranker-v2-base-multilingual: inference latency**

| statistic | library | per_request_ms | per_document_ms |
| --- | --- | --- | --- |
| min | cross-encode-rs | 74.51 | 2.48 |
| min | fastembed | 94.43 | 3.15 |
| min | sentence-transformers | 93.83 | 3.13 |
| p50 | cross-encode-rs | 131.00 | 4.40 |
| p50 | fastembed | 188.53 | 6.34 |
| p50 | sentence-transformers | 190.79 | 6.42 |
| p90 | cross-encode-rs | 180.71 | 6.09 |
| p90 | fastembed | 249.93 | 8.40 |
| p90 | sentence-transformers | 254.38 | 8.54 |
| p99 | cross-encode-rs | 250.75 | 8.36 |
| p99 | fastembed | 340.43 | 11.35 |
| p99 | sentence-transformers | 336.39 | 11.21 |
| max | cross-encode-rs | 581.39 | 19.38 |
| max | fastembed | 480.12 | 16.00 |
| max | sentence-transformers | 584.86 | 19.50 |

_Per request: On the larger Jina model, cross-encode-rs is 1.3 to 1.5x faster than both Python libraries from min to p99 (131 ms versus 189 ms and 191 ms at p50). Only the single slowest request goes to fastembed: 480 ms, against 581 ms for cross-encode-rs and 585 ms for sentence-transformers._

_Per document: cross-encode-rs takes 4.40 ms per document at p50 and 8.36 ms at p99, against 6.34 ms and 11.35 ms for fastembed, and 6.42 ms and 11.21 ms for sentence-transformers._


### Model Load Time

Latency per request is only half of the picture. For cold starts, autoscaling, and serverless deployments, how long a library takes to load a model matters just as much. To compare the libraries rather than the Python interpreter, we timed only the step that creates the ONNX inference session, from inside each process: `init_model()` in `cross-encode-rs`, and the same session setup in `fastembed`. Starting Python, importing libraries, and loading the tokenizer are excluded on both sides. Each library got 11 warmup runs and 41 measured runs per model ([script](https://github.com/AstraBert/cross-encode-rs/blob/main/scripts/model-load-time-bench.sh)).

- **On MiniLM, `fastembed` is slightly faster**, by about 4 ms (12%).
- **On Jina, the two are effectively tied**, 1% apart on average. The higher `cross-encode-rs` p99 comes from a single slow run.

Both libraries spend this time in the same `onnxruntime` call with the same graph optimization level, so there is little room for the language to make a difference. The chart that follows has the exact numbers for each model.

When we timed the whole process instead, `fastembed` took about 134 ms longer than `cross-encode-rs` on both models, even though both loaded the tokenizer and the model. That gap comes from starting the Python interpreter (through uv) and importing the library. It says nothing against fastembed itself, but it is overhead a Python service pays on every cold start.

**Model load time: ONNX session creation only**

| statistic | library | minilm_ms | jina_ms |
| --- | --- | --- | --- |
| mean | cross-encode-rs | 37.6 | 336.6 |
| mean | fastembed | 33.5 | 333.1 |
| p50 | cross-encode-rs | 37.6 | 336.1 |
| p50 | fastembed | 33.5 | 333.1 |
| p90 | cross-encode-rs | 38.0 | 339.5 |
| p90 | fastembed | 33.8 | 337.1 |
| p99 | cross-encode-rs | 38.4 | 395.2 |
| p99 | fastembed | 34.8 | 338.7 |

_MiniLM: fastembed creates the MiniLM session in 33.5 ms on average, 4.1 ms faster than cross-encode-rs at 37.6 ms._

_Jina: On Jina, the two are 1% apart on average: 336.6 ms for cross-encode-rs versus 333.1 ms for fastembed. The cross-encode-rs p99 of 395.2 ms is a single slow run out of 41._


## Conclusion

The headline result is less "Rust beats Python" and more that the runtime underneath your inference call matters more than the language wrapping it. On the small MiniLM model, all three libraries land within 15% of each other at the median, because they all hand the real work to the same `onnxruntime` engine. The one large gap we measured came from configuration, not language: `sentence-transformers` was several times slower until we pinned it to the CPU execution provider. On the larger Jina model, `cross-encode-rs` pulls ahead of both Python libraries by about 1.4x, so once each request carries more work, how a library drives the runtime starts to count as well.

That said, Rust still earned its keep here. A single binary with no Python runtime, no GIL, and a batching worker that we control down to the thread made it straightforward to build a server that stays fast under concurrent load, and to reason precisely about where every millisecond goes. Model loading is not where it wins, though: with Python startup out of the picture, both libraries load a model in about the same time, for the same reason their inference is close. If you are deploying a cross-encoder as its own service rather than inside a larger Python pipeline, that operational simplicity is worth as much as the raw latency numbers.

[`cross-encode-rs`](https://github.com/AstraBert/cross-encode-rs) is open source. If you are building reranking pipelines and want an ONNX-backed cross-encoder server without a Python dependency, give it a try, and tell us where it breaks.
