unsloth studio vs strata engine for qwen3.8-flash-next

#80
by mayankiit04 - opened

Hi Unsloth team,

It seems strata engine has some optimization which is allowing it to run qwen3.8-flash-next on my hardware to run 2x faster on decode and 3x faster on prompt-prefill ..... can we get that same optimization into unsloth studio ??

Hi Unsloth team,

It seems strata engine has some optimization which is allowing it to run qwen3.8-flash-next on my hardware to run 2x faster on decode and 3x faster on prompt-prefill ..... can we get that same optimization into unsloth studio ??

Nope. That'd be a feat of an engineering and a half as unsloth is built to support a lot of models, strata is tailored specifically for one architecture.

i wish if this can be extended for all MoE .... llama.cpp currently has no means to getting higher prob experts onto vram... we are restricted by only -ngl -ncmoe ... i remember back then you had made some changes for ngram streaming... getting all this in one place either under llama.cpp or strata engine will help for more MoE models runtime.

i wish if this can be extended for all MoE .... llama.cpp currently has no means to getting higher prob experts onto vram... we are restricted by only -ngl -ncmoe ... i remember back then you had made some changes for ngram streaming... getting all this in one place either under llama.cpp or strata engine will help for more MoE models runtime.

Maybe one day, but not soon.

Sign up or log in to comment