Text Generation
MLX
Safetensors
English
Chinese
kimi_k2
quantized
kimi
deepseek-v3
Mixture of Experts
instruction-following
4-bit precision
apple-silicon
conversational
custom_code
Instructions to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("richardyoung/Kimi-K2-Instruct-0905-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default richardyoung/Kimi-K2-Instruct-0905-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use richardyoung/Kimi-K2-Instruct-0905-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "richardyoung/Kimi-K2-Instruct-0905-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -52,7 +52,40 @@ This is a **high-performance 4-bit quantized version** of Kimi K2 Instruct, opti
|
|
| 52 |
|
| 53 |
## 🎯 Quick Start
|
| 54 |
|
| 55 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
|
| 57 |
```bash
|
| 58 |
pip install mlx-lm
|
|
|
|
| 52 |
|
| 53 |
## 🎯 Quick Start
|
| 54 |
|
| 55 |
+
#
|
| 56 |
+
|
| 57 |
+
## Hardware Requirements
|
| 58 |
+
|
| 59 |
+
Kimi-K2 is a massive 671B parameter MoE model. Choose your quantization based on available unified memory:
|
| 60 |
+
|
| 61 |
+
| Quantization | Model Size | Min RAM | Quality |
|
| 62 |
+
|:------------:|:----------:|:-------:|:--------|
|
| 63 |
+
| **2-bit** | ~84 GB | 96 GB | Acceptable - some quality loss |
|
| 64 |
+
| **3-bit** | ~126 GB | 128 GB | Good - recommended minimum |
|
| 65 |
+
| **4-bit** | ~168 GB | 192 GB | Very Good - best quality/size balance |
|
| 66 |
+
| **5-bit** | ~210 GB | 256 GB | Excellent |
|
| 67 |
+
| **6-bit** | ~252 GB | 288 GB | Near original |
|
| 68 |
+
| **8-bit** | ~336 GB | 384 GB | Original quality |
|
| 69 |
+
|
| 70 |
+
### Recommended Configurations
|
| 71 |
+
|
| 72 |
+
| Mac Model | Max RAM | Recommended Quantization |
|
| 73 |
+
|:----------|:-------:|:-------------------------|
|
| 74 |
+
| Mac Studio M2 Ultra | 192 GB | 4-bit |
|
| 75 |
+
| Mac Studio M4 Ultra | 512 GB | 8-bit |
|
| 76 |
+
| Mac Pro M2 Ultra | 192 GB | 4-bit |
|
| 77 |
+
| MacBook Pro M3 Max | 128 GB | 3-bit |
|
| 78 |
+
| MacBook Pro M4 Max | 128 GB | 3-bit |
|
| 79 |
+
|
| 80 |
+
### Performance Notes
|
| 81 |
+
|
| 82 |
+
- **Inference Speed**: Expect ~5-15 tokens/sec depending on quantization and hardware
|
| 83 |
+
- **First Token Latency**: 10-30 seconds for model loading
|
| 84 |
+
- **Context Window**: Full 128K context supported
|
| 85 |
+
- **Active Parameters**: Only ~37B parameters active per token (MoE architecture)
|
| 86 |
+
|
| 87 |
+
|
| 88 |
+
## Installation
|
| 89 |
|
| 90 |
```bash
|
| 91 |
pip install mlx-lm
|