AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head and vision tower are preserved as BF16 sidecars when present.

AXQuant checkpoint Tier 1 certified on df-macbookpro-m5 for Hub commit f44a9eeebec0c488d0f42201c8763db770a1c0a8. Product class is 5p6bpw (mixed AXQ; marketing name 4-bit). Size vs uniform-4 is 1.207× under the class 1.25 budget; agent-coding quality retention 0.993, general 1.000. Certificate.

MTP acceleration Tier 2 certified (scoped). On df-macbookpro-m5 with AX Engine 6.14.0, greedy MTP-off/on streams match and decode-heavy profiles clear ≥1.20× / ≥1.10× (agent-coding 1.301× / 1.104×, long-form general 1.223× / 1.249×). Tier 2 certificate.

Product default remains direct fallback. Short-answer chat is not a universal speed claim. Formal route: Qwen linear MTP exact + certification-candidate opt-in.

Model details

Property Value
Base model Qwen/Qwen3.6-27B
Source revision 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Product family qwen3.6
Source architecture Qwen3_5ForConditionalGeneration (dense); text path optimized
Main-model parameters 27.36B logical parameters
Quantizer AXQuant 1.2.0
Hub budget class 4bit
Artifact edition v2
AXQuant base precision class 5p6bpw
Planned storage-adjusted BPW 5.5800
Measured main-model BPW 5.4183
Measured total BPW, including MTP 5.5801
Safetensors weight size 19.38 GB
Approximate complete download 19.40 GB
Configured maximum context 262,144 tokens; practical limits depend on unified memory
MLX-LM compatibility Standard text inference, compatibility level B
AX Engine native execution Tier 1 safe default direct route; scoped Tier 2 MTP certified (opt-in formal contract)
MTP present True
Vision sidecar present True

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models.

Sibling Intended trade-off
4bit sibling Lower-storage AXQ budget; check its exact BPW
6bit sibling Higher average precision near the 6-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP --local-dir ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Allow at least 19.40 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine status

AX Engine 6.14.0 on df-macbookpro-m5 loads this checkpoint for formal certification. Product default remains direct fallback (safe Tier 1). Scoped Tier 2 MTP requires the formal Qwen linear MTP exact / certification-candidate contract (see certificate). MLX-LM remains the standard text inference path and does not by itself establish MTP acceleration.

Use the packaged Qwen MTP head with oMLX or MTPLX

Download the complete repository to a writable local directory. In oMLX 0.6.3rc2 or newer, add that directory, open Model Settings, choose Import MTP side-car, and then enable Lightning MTP. The import changes only the local copy so the sidecar tensors become visible through the checkpoint index. This is a text-path compatibility result; it does not certify VLM loading or vision quality.

MTPLX can consume the packaged sidecar directly:

mtplx quickstart \
  --model ./AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP \
  --profile stable \
  --depth 1 \
  --reasoning off

mtplx_runtime.json declares the canonical qwen3-next-mtp execution contract. This enables strict runtime discovery; it does not establish oMLX or MTPLX exactness or speed certification.

Quantization layout

Main-weight precision Parameters Share
4bit 24.35B 87.65%
8bit 1.27B 4.58%
bf16 2.16B 7.77%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32, 64.
  • MTP sidecar: 15 tensors, 424.70M parameters, 0.85 GB, BF16.
  • Vision sidecar: 333 tensors, 460.73M parameters, 0.92 GB, BF16.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

Check Status
Planning evidence architecture_prior (Tier 1 still uses measured quality/size on host)
Quality vs uniform-4 Certified: agent-coding retention 0.992647; general 1.0
Size vs uniform-4 1.206999 under 5p6bpw max ratio 1.25
MTP exactness + speed (decode-heavy) Tier 2 certified (scoped) — see certificate
Vision-language quality Not claimed; vision tensors preserved at BF16
Full M0–M8 flagship campaign Separate process; not implied by these certificates
Certificates Tier 1 · Tier 2 · index

Modalities (capability-gated)

Text checkpoint Tier 1 does not imply vision or audio quality. Vision present=true on a pack is not a quality pass.

Modality Claim Supported Reason
Vision present-not-certified true vision present sidecar=['vision.safetensors']; mlx-vlm smoke failed on df-macstudio-m2 (see evidence). Text Tier 1 unchanged. Evidence: /Users/akiralam/code/axquant/docs/certifications/evidence/modality-recert-macstudio-m2/results/qwen36-27b-axq4-mtp.json
Audio not-applicable false audio not supported (no tower config and no sidecar weights)

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP acceleration is opt-in under the formal exact contract; default remains direct fallback.

  • Vision weights are byte-preserved at BF16, but this release does not claim validated VLM quality.

  • The configured context window can require substantially more memory as the KV cache grows.

  • AX Engine Tier 1 default is direct fallback; Tier 2 is scoped formal-route only.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. Parallel OptiQ repositories use a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the Qwen/Qwen3.6-27B model card for license terms, model limitations, and responsible-use guidance.

Downloads last month
293
Safetensors
Model size
5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP

Base model

Qwen/Qwen3.6-27B
Quantized
(715)
this model

Collections including AutomatosX/AX-Qwen3.6-27B-MLX-AXQ-4bit-MTP