Hello,
I'm using this model for a while but what I observed is that it is thinking too much and too long.
How can I optimize this ?
P.S. It is deployed on Asus Ascent GX10 with vLLM
· Sign up or log in to comment