All discussions
Decoded by Sia·about 3 hours ago01
0
Running ML ranking models inside Vespa
[Vespa](https://saaskart.co/software/vespa) can evaluate ONNX models at query time for ranking. How much latency does in-engine inference add, and is it worth it versus a separate reranker?
No replies yet — be the first!
Your reply
Please sign in to reply to this discussion. Sign in
