MiMo-V2.6-Pro-EXL3

Experimental sequential quantization checkpoint. All 69 MoE layers are complete (layer 0 is dense), and it can be served with the custom loader and kernels at https://github.com/jarrelscy/mimo-v2.6-pro-exl3-sm120. It is not an off-the-shelf EXL3 model directory, so exllamav3/TabbyAPI, vLLM and SGLang cannot load it directly.

Serving

Tested on 4x RTX PRO 6000 Blackwell (SM120, 96 GB, PCIe P2P, no NVLink): TP4, ~72 GiB per GPU, 11.8 ms/step, ~83 tok/s single-stream decode. Needs the CUDA 12.8 toolkit (the extensions are JIT-built on first start), Python 3.12 and ~283 GB of disk.

hf download jarrelscy/MiMo-V2.6-Pro-EXL3 --local-dir /path/to/MiMo-V2.6-Pro-EXL3   # experts.tar is read in place

git clone https://github.com/jarrelscy/mimo-v2.6-pro-exl3-sm120 && cd mimo-v2.6-pro-exl3-sm120
python3.12 -m venv venv && venv/bin/pip install -r requirements.txt
export MIMO_VENV=$PWD/venv CUDA_HOME=/usr/local/cuda-12.8 MIMO_EXL3_DIR=/path/to/MiMo-V2.6-Pro-EXL3
./run_server.sh            # OpenAI-compatible server on :8003

curl localhost:8003/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model": "local", "messages": [{"role": "user", "content": "The capital of France?"}], "max_tokens": 256}'

The server provides /v1/completions and /v1/chat/completions (streaming, reasoning_content, tool calls). It serves one request at a time with a 32K context by default (MIMO_LMAX). Text only: vision and audio are not wired up. MTP speculative decoding (MIMO_SPEC=K) is experimental and off by default. See the GitHub README for the options.

Checkpoint

Completed full-corpus joint-PV layers: [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69]. Each completed layer's run traversed all 18,006,461 training tokens. Validation alone selects the retained checkpoint, which may precede the end of the corpus pass or retain the initial fit when it is better. The selected step and corpus coverage are recorded in reports.

Cold experts use mixed-rate EXL3 trellis codes at <=2 actual packed bits per weight. Initial-fit provenance is recorded per layer. Sequential repair from layer 23 fits candidates on all training examples routed to each cold expert and retains existing weights if same-input validation does not improve. Layers 1–22 retain their previous weights. Coverage and selection decisions are recorded in reports; subsequent full-corpus PV jointly optimizes all input/output scales against the combined routed layer output while holding trellis codes fixed. No neurons are removed. The exact original 1,325 hot experts (5.00075% globally) retain their NVFP4 weights and assignment; this is not a separate 5% allocation per layer.

Later layers use inputs propagated through earlier accepted EXL3 layers. Validation and audit each use 16,384 held-out tokens. reports/layer_errors.csv records EXL3 validation and audit reconstruction errors. ARVQ comparisons are disabled; previously published comparison reports are historical. Relative L2 errors are layer reconstruction metrics, not whole-model accuracy, perplexity, or KL.

Cold weights use standard EXL3 codebooks, not the experimental FP4-component codebooks. Evaluation decodes weights to BF16. These custom mixed-rate per-expert artifacts need the loader linked above. The backbone (attention, router, norms, embeddings, lm_head, MTP) is in backbone/ and the NVFP4 hot experts are in hot/.

Layers 26 onward store exact expert binaries and receipts in layers/layer_NNNNN/experts.tar to stay within repository file limits. The serving loader reads the archive in place; to use decode.py, extract it with tar -xf experts.tar; member paths preserve the original layers/layer_NNNNN/expert_NNNNN/selected.bin layout. Per-layer manifests record both archive and individual expert SHA-256 hashes. Earlier individual binary paths remain available.

Use decode.py with the official EXL3 1.5.1 runtime and a compatible CUDA/PyTorch wheel to reconstruct an expert:

python decode.py layers/layer_00001/expert_00000/selected.bin expert_00000.pt

Source: XiaomiMiMo/MiMo-V2.6-Pro-RL, revision 73875d00b30a89ef8cc353a0b60b0e9f9561952d. See sequential_manifest.json, per-layer manifests, and receipts for status, provenance, and SHA-256 hashes. The earlier FP4-named research repository is separate.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jarrelscy/MiMo-V2.6-Pro-EXL3

Finetuned
(3)
this model