danielhanchen commited on
Commit
28fadd5
·
verified ·
1 Parent(s): affab4c

Add llama.cpp PR build instructions

Browse files
Files changed (1) hide show
  1. README.md +24 -0
README.md CHANGED
@@ -57,6 +57,30 @@ Note: MiniMax Sparse Attention is not supported yet, so inference falls back to
57
 
58
  # MiniMax-M3
59
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
  MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
61
 
62
  **Highlights:**
 
57
 
58
  # MiniMax-M3
59
 
60
+ ## Run MiniMax-M3 in llama.cpp
61
+
62
+ MiniMax-M3 support in llama.cpp is preliminary and not yet in a released build. To run these GGUFs, build llama.cpp from [PR #24523](https://github.com/ggml-org/llama.cpp/pull/24523):
63
+
64
+ ```bash
65
+ git clone https://github.com/ggml-org/llama.cpp
66
+ cd llama.cpp
67
+ git fetch origin pull/24523/head:minimax-m3
68
+ git checkout minimax-m3
69
+ cmake -B build -DGGML_CUDA=ON
70
+ cmake --build build --config Release -j --target llama-cli llama-server
71
+ ```
72
+
73
+ Then run a quant. The model is large (~428B params), so offload across GPUs with `-ngl 99` or keep the weights in CPU RAM:
74
+
75
+ ```bash
76
+ ./build/bin/llama-cli \
77
+ -hf unsloth/MiniMax-M3-GGUF:UD-Q4_K_XL \
78
+ --jinja -ngl 99 --ctx-size 8192 \
79
+ -p "Hello, who are you?"
80
+ ```
81
+
82
+ Note: MiniMax Sparse Attention is not supported yet, so inference falls back to dense attention.
83
+
84
  MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
85
 
86
  **Highlights:**