FrogNano-4B-2609 -- Pollard

Pollard shrank this model: 8.65 GB (f16) -> 1.42 GB -- 84% smaller, 6.1x down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 8.65 GB
Q8_0 ~4.59 GB
Q6_K 3.56 GB
Q4_K_M ~2.51 GB
PollardMix (this repo's IQ2_KT) 1.42 GB

Pollard builds of microsoft/FrogNano-4B-2609 made with Pollard Weights -- a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF -- runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio, except where noted. IQ2_KT needs ik_llama.cpp: its allocation puts ik_llama-only atoms on the tensors it protects. The rest run anywhere.

Model details

Parameter count ~4.3B
Architecture qwen3_5
Input support text, image
imatrix yes -- see calibration
Perplexity measured yes -- WikiText-2, 64 × 2048-token chunks, vs f16 (PPL 9.338)

Which file should I choose?

Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:

  • ~5.6 GB RAM / VRAM -> Q6_K (3.56 GB).
  • ~4.5 GB RAM / VRAM -> IQ4_XS (2.48 GB).
  • ~4.1 GB RAM / VRAM -> IQ3_S (2.06 GB).
  • ~3.4 GB RAM / VRAM -> IQ2_KT (1.42 GB). (ik_llama.cpp)

Available files

file PPL size Mean KLD runs in notes
FrogNano-4B-2609-Pollard-IQ2_KT.gguf 13.51 1.42 GB 0.614 ik_llama smallest · top-1 68.0%
FrogNano-4B-2609-Pollard-IQ3_S.gguf 9.59 2.06 GB 0.248 any llama.cpp top-1 80.7%
FrogNano-4B-2609-Pollard-IQ4_XS.gguf 9.08 2.48 GB 0.185 any llama.cpp top-1 83.8%
FrogNano-4B-2609-Pollard-Q6_K.gguf 9.44 3.56 GB 0.019 any llama.cpp largest · top-1 96.0% · near-lossless

Measured on WikiText-2 (64 chunks × 2048 tokens) against the f16 original (PPL 9.338): Mean KLD = average divergence of each rung's next-token distribution from f16 (lower is closer); top-1 = how often the rung's most likely next token matches f16's. FrogNano is an instruct/thinking model, so raw-text perplexity reads higher than on chat text; compare rungs by KLD.

Coherence gate

Every rung is checked by Pollard's coherence gate before it ships. The IQ2_KT flagship (1.42 GB, ik_llama.cpp) passes all three gate prompts (factual, explanation, code) and stops on its own end-of-turn token. Recommended sampling, from the gate's sweep:

--temp 0.7 --repeat-penalty 1.15 --repeat-last-n 256 --top-k 40 --top-p 0.9

The allocation behind these rungs is measured per layer on FrogNano itself (pollard-probe, 32 layers). Perplexity and mean KLD against f16 for every rung are in the table above: Q6_K is near-lossless (mean KLD 0.019, 96% top-1 agreement with f16).

Multimodal

Vision needs the projector shipped alongside: mmproj-FrogNano-4B-2609-BF16.gguf -- download it too and pass it with --mmproj. It is not quantized; it is small and the text ladder is where the size lives.

llama-server -m FrogNano-4B-2609-Pollard-IQ2_KT.gguf --mmproj mmproj-FrogNano-4B-2609-BF16.gguf -ngl 99

Download a specific file

pip install -U "huggingface_hub[cli]"
hf download PollardWeights/FrogNano-4B-2609-Pollard \
  --include "FrogNano-4B-2609-Pollard-IQ2_KT.gguf" --local-dir ./

How to run

IQ2_KT is built on ik_llama-only atoms, so it runs with ik_llama.cpp:

llama-cli    -m FrogNano-4B-2609-Pollard-IQ2_KT.gguf -ngl 99 -p "Explain why the sky is blue."
llama-server -m FrogNano-4B-2609-Pollard-IQ2_KT.gguf -ngl 99

For stock llama.cpp, Ollama or LM Studio, use Q6_K instead:

llama-server -hf PollardWeights/FrogNano-4B-2609-Pollard:Q6_K
llama-cli    -m FrogNano-4B-2609-Pollard-Q6_K.gguf -ngl 99 -p "Explain why the sky is blue."

imatrix (calibration)

The importance matrix (FrogNano-4B-2609-Pollard.imatrix, included) was computed on Calib 3.0 multi-domain calibration.

ARM / AVX

llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines -- no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.

Errata

  • IQ2_KT carries ik_llama-only atoms and needs ik_llama.cpp to run; stock llama.cpp rejects any ggml type above 42 outright. Checked with pollard-ggufcheck, from the files' tensor types rather than their names.
  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Credits & license

Built with Pollard Weights -- frontier models, small hardware, no compromise.

Downloads last month
805
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PollardWeights/FrogNano-4B-2609-Pollard

Quantized
(17)
this model