clef -- Pollard

Pollard shrank this model: 54.05 GB (f16) -> 9.44 GB -- 83% smaller, 5.7x down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 54.05 GB
Q8_0 ~28.65 GB
Q6_K 22.19 GB
Q4_K_M ~15.67 GB
PollardMix (this repo's IQ2_XXS) 9.44 GB

Pollard builds of Cloudflare/clef made with Pollard Weights -- a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF, but you need a recent llama.cpp. This model's architecture (clef) is implemented upstream, so any build new enough to carry it runs these files -- llama.cpp itself, and Ollama or LM Studio once they ship a runtime with it. An older build will refuse them with unknown model architecture. The quants are ordinary K-quants.

Model details

Parameter count ~27.0B
Architecture clef
Input support text, image
imatrix yes -- see calibration
Measured decision fidelity vs Q8_0 through /v1/systemone -- table below (perplexity does not apply: a decision model answers with probabilities, not text)

Which file should I choose?

Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:

  • ~24.2 GB RAM / VRAM -> Q6_K (22.19 GB). (needs a current llama.cpp)
  • ~18.2 GB RAM / VRAM -> IQ4_XS (16.19 GB). (needs a current llama.cpp)
  • ~15.3 GB RAM / VRAM -> IQ3_S (13.25 GB). (needs a current llama.cpp)
  • ~11.4 GB RAM / VRAM -> IQ2_XXS (9.44 GB). (needs a current llama.cpp)

Available files

file size agrees with Q8_0 option KL notes
clef-Pollard-IQ2_XXS.gguf 9.44 GB 100% 0.0038
clef-Pollard-IQ3_S.gguf 13.25 GB 100% 0.00049
clef-Pollard-IQ4_XS.gguf 16.19 GB 100% 9.5e-05
clef-Pollard-Q6_K.gguf 22.19 GB 100% 2.9e-05

Decision fidelity

This is a decision model: it answers typed questions (choice / score / yes-no) with a probability per option, so it is measured on its decisions, not on text. Each rung and the Q8_0 answered the same 20 typed questions through llama-server's /v1/systemone, read the same way (pollard-decision).

file size agrees with Q8_0 option KL vs Q8_0 mean prob. drift accuracy
clef-Pollard-Q6_K.gguf 22.19 GB 100% 2.9e-05 0.0003 100%
clef-Pollard-IQ4_XS.gguf 16.19 GB 100% 9.5e-05 0.0007 100%
clef-Pollard-IQ3_S.gguf 13.25 GB 100% 0.00049 0.0014 100%
clef-Pollard-IQ2_XXS.gguf 9.44 GB 100% 0.0038 0.0066 100%

Agrees = the rung picks the same option as the Q8_0. Option KL = how far its option probabilities moved from the Q8_0's (0 = identical); it is the number that separates the rungs.

Build notes

  • Measured allocation. The sensitivity profile was derived from the imatrix and the f16 weights one tensor at a time (pollard-probe --from-imatrix), covering all 64 layers, including the 48 Gated DeltaNet linear-attention layers. Profile included: clef-Pollard.sensitivity.json. imatrix over the full Calib 3.0 corpus (1937 chunks), computed on a Q8_0 copy: clef-Pollard.imatrix.
  • Reference. The decision board compares each rung with Q8_0: the 54 GB f16 does not fit the build machine's RAM + VRAM for serving.
  • Decision head precision. The decision head (dec.*) is held at Q6_K in every rung. Rebuilding the IQ2_XXS rung with the head at BF16 (same plan otherwise) gave option KL 0.00377 vs 0.00385 with the Q6_K head: the head at Q6_K does not collapse the probabilities, and the drift that remains comes from the body.
  • Images. The projector carries clip.vision.image_min_pixels = 65536 from the model's processor, so small images get the same number of tokens as the original processor on current llama.cpp.

Multimodal

Vision needs the projector shipped alongside: mmproj-clef-BF16.gguf -- download it too and pass it with --mmproj. It is not quantized; it is small and the text ladder is where the size lives.

llama-server -m clef-Pollard-IQ2_XXS.gguf --mmproj mmproj-clef-BF16.gguf -ngl 99

Download a specific file

pip install -U "huggingface_hub[cli]"
hf download PollardWeights/clef-Pollard \
  --include "clef-Pollard-IQ2_XXS.gguf" --local-dir ./

How to run

This model's architecture (clef) needs a llama.cpp new enough to carry it. A decision model is served on /v1/systemone (no chat or completions):

llama-server -m clef-Pollard-IQ2_XXS.gguf -ngl 99 --port 8080
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
  "questions": {
    "route":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": null, "shipping": null, "technical": null}},
    "angry":   {"type": "noul",   "instructions": "Is the customer angry?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["can wait", "this week", "today", "right now"]}
  }
}'

Each answer carries the probability of every option. Some decision models evaluate the whole prompt in one batch, so a long state may need a larger --ubatch-size.

imatrix (calibration)

The importance matrix (clef-Pollard.imatrix, included) was computed on a mixed-domain corpus so the matrix sees every register the model serves.

ARM / AVX

llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines -- no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.

Errata

  • general.architecture is clef, which upstream llama.cpp added recently. A build older than that support refuses these files with unknown model architecture -- update llama.cpp rather than looking for a different quant. Checked with pollard-ggufcheck, which reads the architecture and the tensor types out of the header and asks upstream what it implements.
  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Credits & license

  • Base model: Cloudflare/clef
  • Quantization tooling: llama.cpp (ggml-org)
  • Method + tooling: Pollard Weights -- measure first, no claim before a number.
  • License: apache-2.0, inherited from the base model.

Built with Pollard Weights -- frontier models, small hardware, no compromise.

Downloads last month
45
GGUF
Model size
27B params
Architecture
clef
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/clef-Pollard

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(35)
this model