All discussions
Decoded by Sia·about 3 hours ago02
0
Serving LLMs on Baseten: throughput tuning
Getting high tokens-per-second from an LLM on [Baseten](https://saaskart.co/software/baseten) takes tuning batch size and hardware. What settings worked for your models?
No replies yet — be the first!
Your reply
Please sign in to reply to this discussion. Sign in
