Siddhartha Venkatayogi

Relevance Ranking with Jev and Qwev (Qwen)


This weekend, I experimented with relevance ranking using Jev by TypeSafe AI. Here are my thoughts!


The Approach

I framed the task into four classification probabilities: Exact, Substitute, Complement, and Irrelevant. Then I used the probabilities to calculate a relevance score.

I also adapted Together AI’s Tev methodology to train a separate Qwen3.5-4B model on ~20k examples, using 1× NVIDIA RTX 6000 Ada with 48 GB VRAM. Instead of generating an answer token, it uses a four-class head and returns probabilities directly (no autoregressive token generation and decode step!!)

Ranking Results

Results using Amazon’s shopping-query dataset. NDCG@10 is higher-is-better, and latency is the median per query.

Model NDCG@10 Median latency per query Hardware / serving
Trained Qwen .809 1,463 ms 1× RTX 6000 Ada 48 GB, via Runpod
Jev .807 197 ms TypeSafe network API
BGE reranker .770 44 ms 1× RTX 6000 Ada 48 GB, via Runpod
Untuned Qwen .745 1,442 ms 1× RTX 6000 Ada 48 GB, via Runpod
MiniLM reranker .740 19 ms 1× RTX 6000 Ada 48 GB, via Runpod
BGE embeddings .739 7 ms 1× RTX 6000 Ada 48 GB, via Runpod
BM25 .703 0.67 ms Runpod host x86 CPU; exact CPU model not recorded



Classification Accuracy

Classification accuracy on the same 20,216 test pairs. Four-class accuracy uses Exact / Substitute / Complement / Irrelevant. For binary accuracy, Exact + Substitute + Complement are considered Relevant.

Model Four-class accuracy Binary accuracy
Trained Qwen 60.61% 85.47%
Jev 56.32% 79.64%
Untuned Qwen 50.66% 83.34%

The predict-always-relevant baseline is 83.42% for this dataset.

My Takeaways

The simply trained Qwen didn’t get a significant NDCG advantage over Jev, which is fully zero-shot! Jev is pretty impressive, but also I have no idea about parameter counts/comparisons.

I do think I could’ve had a significantly stronger training protocol for Qwen. I think a teacher-student protocol and also training my own encoder would have better results. I’m probably underutilizing the current parameters now.

Regarding the speed: without the decode step, prefill dominates processing and is the primary bottleneck (mostly compute-bound). BGE reranking was ~33× faster on the same GPU, while giving up some ranking quality.

Qwen used reference kernels, so this didn’t fully optimize serving/inference. First things to add are maybe better linear attention kernels and prefix caching.

This entire project cost about $5.36 (Jev was ~$0.50, the rest was Runpod GPU costs).



Make a classifier or just use Jev?

Jev could really help with relevance + ranking without having to train and host large foundational models yourself. If you’ve been wanting to implement some candidate ranking for any kind of recommendation features on your platform, but haven’t had the bandwidth to make it good, Jev is a promising option, while being relatively cheap 😸