NInfer models

Quantized artifacts for NInfer Ext, a from-scratch C++/CUDA inference engine for maximum single-GPU performance.

These artifacts only work with giveen/ninfer-ext. They are not Transformers checkpoints, not GGUF, not safetensors weights, and cannot be loaded by transformers, vLLM, llama.cpp, or exllamav3. .ninfer is NInfer's own artifact format, and the EXL3 trellis and NVFP4 layouts are decoded by kernels that live in that repository. Loading them anywhere else will fail.

Contents

Model Artifact Size Bundled
Qwen3.8-27B EXL3 4.0 bpw qwen3_8_27b_exl3_4bpw.ninfer 15.68 GiB Text + MTP + Vision
Qwen3.8-27B EXL3 3.5 bpw qwen3_8_27b_exl3_3p5bpw.ninfer 14.24 GiB Text + MTP + Vision
Qwen3.8-Flash-Next NVFP4 qwen3.8-flash-next/ 119 GB Text + MTP + Vision
Qwen3.8-Flash-Next EXL3 4.0 bpw qwen3.8-flash-next-exl3-4bpw/ 95.82 GB Text + MTP
Qwen3.8-Flash-Next EXL3 3.5 bpw qwen3.8-flash-next-exl3-3p5bpw/ 88.11 GB Text + MTP
Qwen3.6-35B-A3B NVFP4 qwen3.6-35b-a3b-nvfp4/ 20.26 GiB Text + Vision + MTP

Every artifact ships with its .conversion.json provenance and a SHA256SUMS; the folders additionally carry a README with serve and convert commands. Per-model details are below.

Qwen3.8-27B EXL3

Model

Qwen3.8-27B (Qwen3_5ForCausalLM): 64 layers (48 GDN linear-attention, 16 full attention), hidden 5120, intermediate 17408, vocab 248,320, plus a separate MTP layer and a Vision tower.

Quantization

NInfer's native EXL3 format (exl3_mul1 + trellis_t16_v1) β€” a three-instruction trellis codebook over 16x16 tiles, with the input and output Hadamard rotations folded into the kernels. Weights are produced by NInfer's own C++/CUDA quantizer (ninfer-quantize) from full-precision source tensors; no exllamav3 checkpoint is imported.

Per-tensor rates (half bits, i.e. X.5 bpw, are first-class trellis rates):

Scope 4.0 bpw artifact 3.5 bpw artifact
MLP and GDN projections 4.0 bpw 3.5 bpw
Attention projections (the -hq promotion) 5.0 bpw 4.5 bpw
Vocabulary head 6.0 bpw 6.0 bpw
MTP layer 4.0 bpw (5.0 attention), calibrated 3.5 bpw (4.5 attention), calibrated
Vision tower groupwise Q6/Q8/Q4/Q5 (not EXL3 β€” its MLP intermediate is not 128-aligned) same

The MTP layer is calibrated from the final hidden states and next-token embeddings.

Both artifacts are the same model at two points on the size curve. 4.0 bpw is the one to lead with: it is the only artifact here that beats both of the engine's other Qwen3.8-27B builds on both perplexity and KL divergence, while being smaller than either (see Quality). 3.5 bpw is the size-optimised tier β€” best PPL per byte, but it does not carry that advantage into divergence.

Quality

Full-corpus perplexity over 261,167 tokens (context/stride 4096/2048, FP8 KV, greedy), and KL divergence against the full-precision model over the same 2,940 positions of a 23.5k-token wikitext slice:

Artifact Size PPL KL(P_BF16 β€– P)
EXL3 4.0 bpw 15.68 GiB 4.2939 0.0332
Groupwise INT4 16.96 GiB 4.3439 0.0429
NVFP4 22.09 GiB 4.3149 0.0510
EXL3 3.5 bpw 14.24 GiB 4.3101 0.0624

The two metrics rank these differently, and both are reported for that reason: 4.0 bpw wins on both, but 3.5 bpw's better perplexity than INT4 and NVFP4 does not survive as divergence. Only the KL ordering should be compared across runs β€” the absolute values move by roughly 2Γ— with the text, and the ordering was reproduced in both halves of the reference.

Performance

Measured on one NVIDIA GeForce RTX 5090, CUDA 13.3, a single request, greedy, 64-256 output tokens, --prefill-chunk 1024:

Regime EXL3 4.0 bpw EXL3 3.5 bpw Q4 NVFP4
Decode, plain 75 tok/s 64 tok/s 83 tok/s 73 tok/s
Decode, MTP K=3 137 tok/s 130 tok/s 134 tok/s 142 tok/s
Decode, MTP K=5 + --lm-head-draft 145 tok/s 131 tok/s 144 tok/s 167 tok/s
Prefill, 0.54k / 7.6k-token prompt 1.97k / 2.33k tok/s 1.72k / 2.11k tok/s 2.42k / 2.91k tok/s 5.28k / 8.60k tok/s

Measuring MTP at K=3 for all four keeps the draft length equal across formats; the K=5 row is each artifact's own best setting, which Q4 and NVFP4 reach with --lm-head-draft. Both EXL3 tiers carry the indexed proposal head that flag needs, so it is available to them too.

4.0 bpw is close to the other native formats on every regime: prefill 1.19–1.23x behind Q4, plain decode within 10%, and at a comparable draft length its MTP matches Q4's exactly, on an artifact 1.3 GiB smaller (15.68 against 16.96 GiB). NVFP4 remains the prefill leader β€” as it is for this engine's other models β€” because its tensor-core contraction needs no per-weight decoding, which a 4-bit trellis does: the contraction issues exactly the same number of MMAs as Q4's, and the difference is the funnel, bit-field extracts and IMAD/DP4A per decoded window that the trellis costs.

3.5 bpw is slower than 4.0 bpw on every regime, on the same kernels. Its weights are 9% smaller, but its odd half-rates take the heavier exl3_windows_half window decode β€” two funnel shifts for the eight windows against one funnel and five bit-field extracts β€” and that costs more than the bytes it saves. Its case is size and perplexity per byte, not speed.

These are single-request spot measurements, not the engine's methodology-conforming performance tables; docs/performance.md records the published coverage and the difference.

Qwen3.8-Flash-Next NVFP4

Qwen3.8-Flash-Next (Qwen4ExpForCausalLM) has about 180B parameters: about 121B are 512 routed experts per layer, and 51B are an n-gram embedding table. This artifact is converted from nvidia/Qwen3.8-Flash-Next-NVFP4 with the qwen3_8_flash_next_nvfp4 recipe, and contains Text, MTP and Vision.

Representation

Weights Stored as Runtime residency
Routed experts (48 Γ— 512, plus the MTP layer) NVFP4, imported codes and scales; MTP re-encoded from block FP8 pinned Host, fetched into a device expert cache
N-gram PLE table (320M Γ— 160) FP8 rows with BF16 multipliers page-cache mapped or streamed from NVMe, gathered on the Host per token
Attention, GDN, hyper-connection, shared expert, PLE projections Q8 device
Token embedding / output head Q8 / Q6 device
Routers, shared-expert gates, norms, small vectors BF16/FP32 direct device

The routed experts run on the W4A4 tensor-core route; each MoE layer resolves its top-10 experts against an LRU device expert cache.

Measured

One RTX 5090 (32 GB, sm_120a), CUDA 13.3, --expert-cache auto, fp8 KV, --spec mtp:

Metric Value
Causal perplexity (ninfer-ppl-1m-v1, quick, fp8 KV) 3.518
Prefill, 262,144-token budget, FP8 KV, --ngram-residency stream, single request 1,090 / 3,666 / 3,498 / 2,908 / 2,218 tok/s at 1.5k / 8k / 64k / 128k / 256k tokens
Decode (greedy, MTP K=3) 81 tok/s
Peak host RSS (--ngram-residency stream) ~65 GiB
Artifact size 119 GB, 4 sharded files

Prefill is fastest at medium prompts: a short prompt is dominated by the fixed per-chunk expert streaming, and a long one by the QSA selection. It also depends on how the n-gram table is read: with it mapped through the page cache (--ngram-residency mapped, the default when host memory allows) prefill is about 1.7x faster than the stream figures above β€” 6,250 tok/s at 8k β€” at the cost of keeping the ~52 GB table in RAM. Long-context prefill improved 1.48x at 64k, 1.87x at 128k and 2.53x at 256k over the previous engine, because the QSA block selection now pools each block's index keys once per select call instead of once per query column (bit-identical).

Requirements

  • One RTX 5090 (sm_120a) and CUDA 13.3.
  • About 70 GB of host RAM (measured ~65 GiB peak RSS) for the pinned experts with --ngram-residency stream, which reads the n-gram table from NVMe. Keeping the ~52 GB table in the page cache (mapped, chosen automatically when memory allows) needs more RAM and is faster once warm.
  • KV storage bf16, int8, fp8, nvfp4 or k8v4; speculative decoding --spec mtp.

Qwen3.8-Flash-Next EXL3

The same model as above with the 48 Γ— 512 routed experts (and the MTP layer's) re-quantized to NInfer's EXL3 exl3_mul1 trellis format at a flat 4.0 or 3.5 bpw, per-expert Hessians from calibration. Dense projections are Q6 (Q8 where K is not a multiple of 128) and the n-gram table is 4-bit row-grouped, so the artifacts are 95.82 GB and 88.11 GB against 119 GB for NVFP4. Experts stay in pinned Host memory behind the device expert cache, as for NVFP4.

Measured

One RTX 5090 (32 GB, sm_120a), CUDA 13.3, i9-285K, --spec mtp --draft-tokens 3, fp8 KV, --expert-cache auto:

Metric EXL3 4.0 bpw EXL3 3.5 bpw
Artifact size 95.82 GB 88.11 GB (-8.0%)
Perplexity, 261k-token mixed text (4096 / 2048, fp8 KV) 3.525 3.535
KL divergence vs the NVFP4 artifact, same text 0.0585 0.0629 (+7.5%)
Decode, 7 mixed chat requests, mean per request, E-cores 172.7 tok/s 186.7 tok/s (+8.1%)
Expert-cache hit rate on that traffic 90.2% 92.0%
Prefill, 8k-token prompt about 6,050 tok/s not measured pinned
  • KL is against the NVFP4 artifact, not BF16; no BF16 baseline was run, and the perplexity text is not the corpus of the NVFP4 figure above, so the two perplexities are not directly comparable.
  • Decode depends strongly on CPU core placement on a hybrid CPU: pinned to the P-cores the same runs gave 121.7 and 133.0 tok/s (about 30% slower). The figures above are pinned to the E-cores (taskset -c 8-23).
  • Agentic and coding quality is not measured. A SlopCodeBench attempt (mini-swe, greedy, no thinking) was inconclusive: 4.0 bpw scored 0.18 weighted test pass rate, and the 3.5 bpw run was stopped after its agent fell into a greedy repetition loop on the first checkpoint. Do not use greedy decoding for agent work.
  • Long agent contexts prefill slowly (near 250 tok/s at about 47k tokens, 44% prompt reuse); not investigated.

Folder READMEs carry the per-artifact details, serve and convert commands.

Qwen3.6-35B-A3B NVFP4

Qwen3.6-35B-A3B (Qwen3_5MoeForCausalLM) is a hybrid full-attention + GatedDeltaNet mixture-of-experts: 40 layers, hidden 2048, 256 experts (top-8, intermediate 512), vocab 248,320, plus a Vision tower and a separate MTP layer. This artifact is converted from nvidia/Qwen3.6-35B-A3B-NVFP4 with the qwen3_6_35b_a3b_nvfp4 recipe, and contains Text, Vision and MTP.

Representation

Weights Stored as
Routed + shared experts (256 experts; gate/up 512Γ—2048, down 2048Γ—512) NVFP4, imported ModelOpt codes, block scales and divisors (W4A4)
Attention, GDN and Vision projections Q8
Token embedding / output head Q8 / Q6
Routers, shared-expert gates, norms, GDN a/b, small vectors BF16 / FP32 direct

The routed and shared experts are the only NVFP4 mechanism; every other projection matches the engine's groupwise-int build for the same model, so the two artifacts differ in exactly one mechanism and are directly comparable.

Measured

One RTX 5090 (32 GB, sm_120a), CUDA 13.4, bf16 KV, spec=none unless noted. Baseline is the engine's Qwen3.6-35B-A3B groupwise-int build (Q4 gate/up, Q5 down, A16 activations):

Metric NVFP4 groupwise-int Delta
Prefill pp2048 27,538 tok/s 19,445 tok/s +41.6%
Prefill pp8192 25,300 tok/s 18,345 tok/s +37.9%
Decode tg128, plain 405 tok/s 404 tok/s ~0%
Decode, MTP K=3 575 tok/s (accept 0.70) 520 tok/s (accept 0.58) +10.5%
Causal perplexity (ninfer-ppl-1m-v1 quick, int8 KV) 4.455 4.365 +2.1%
Artifact size 20.26 GiB 21.22 GiB -4.5%

The prefill gain is the native Blackwell FP4 block-scale MMA: the NVFP4 route runs the routed MoE without the 4-bit β†’ bf16 weight dequantization that caps the groupwise route. The perplexities above use int8 KV, unlike the 27B and Flash-Next tables' fp8 KV, so they are not comparable across models.

Speculative decoding

  • MTP is the recommended backend (--spec mtp --draft-tokens 3).
  • DFlash (v1) runs but is slower than plain decoding on this target, so it is not shipped.
  • DFlash2 for this model is not runnable in the current engine.

Usage

Build the engine from source, then serve an artifact:

git clone https://github.com/giveen/ninfer-ext
cd ninfer-ext && cmake -B build -DCMAKE_BUILD_TYPE=Release \
  -DPython3_EXECUTABLE=$PWD/.venv/bin/python && cmake --build build -j

# Qwen3.8-27B EXL3
build/apps/ninfer-serve qwen3_8_27b_exl3_4bpw.ninfer \
  --port 8099 --spec mtp --draft-tokens 3 --fixed-draft

# Qwen3.8-Flash-Next NVFP4
build/apps/ninfer-serve qwen3.8-flash-next/qwen3_8_flash_next_nvfp4.ninfer \
  --model-id qwen3.8-flash-next --max-context 229376 --kv-capacity 458752 \
  --kv-dtype fp8 --expert-cache auto --ngram-residency stream --spec mtp

# Qwen3.6-35B-A3B NVFP4
build/apps/ninfer-serve qwen3.6-35b-a3b-nvfp4/qwen3_6_35b_a3b_nvfp4.ninfer \
  --model-id qwen3.6-35b-a3b-nvfp4 --max-context 65536 --kv-capacity auto --kv-dtype bf16 \
  --spec mtp --draft-tokens 3

--vision enables image input; the Vision tower loads lazily. See the repository's README.md, docs/cli.md, and docs/serving.md for the full command surface.

License

The quantized weights follow the base model's license. The quantization format, kernels, and tooling are part of giveen/ninfer-ext.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for jabbatheduck/ninfer-ext-models

Finetuned
(355)
this model