MLX version for Qwen3.6 Models?

#7
by cnsiva - opened

Is there a plan to create a MLX version for the Qwen3.6 models?

Owner

image
already done locally, when I get a few free moments to upload to HF.

Ex0bit, I was wondering if there’s any chance we could get the Mlx version published?

Owner

@cnsiva - here you go: https://hf.co/Ex0bit/Qwen3.6-35B-A3B-PRISM-MLX-NVFP4 - Spread the word and enjoy!

Thank you @Ex0bit . I am a big fan of your MYTHOS-26B-A4B-PRISM-PRO

Owner

You're welcome, thank you for the support. Enjoy.

Thank you, Ex0bit.

I tested this model on my M1 Max (64 GB) and achieved approximately 40 t/s in a single session. However, after testing it with two coding agent sessions, the t/s dropped to below 10 for each session.

I typically get 50 to 55 t/s for the Qwen3.5-A3B models (6-bit quantization) and around 30 t/s per session.

Is there any specific configuration required for the oMlx? I used your script to launch the oMlx.

I am testing it from my M5 MAX (64 GB), here is the screenshot

image

Sign up or log in to comment