humanizer: rewrites AI drafts so they read like a person wrote them The humanizer app running these GGUF files locally: draft on the left, rewrite on the right, new wording highlighted

humanizer 12B: GGUF quantizations

GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, trained to keep every number, date, name and quote. The 2-bit, Q3 and Q4_K_M files are quantization-aware trained and distilled from bf16, so they stay closer to the full model than standard quantizations of the same size.

Quick start: the desktop app (Mac, Windows) downloads these files for you; or hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json and see Usage.

Results

Evaluation card: 95% of English rewrites judged human by Originality.ai at its strictest setting (11 of 210 flagged; previous release 26); 376 of 420 English rewrites with no factual problem (previous release 369)

95% judged human by Originality.ai at its strictest setting (210 English drafts, bf16 weights, 2026-10-02): 11 of 210 rewrites were flagged as AI, against 26 for the previous release, and no AI detector was used anywhere in training. On the same 60 drafts, the 2-bit file here had 0 of 60 flagged (Q8_0 7, bf16 4); per-file results are under Quality on the task. 376 of 420 English rewrites came back with no factual problem from a strict LLM judge, measured on the Q8_0 file here (210 drafts, two samples each); where it did find one, more than 9 in 10 fixes are a single word or phrase. Full results: main model card.

Smaller files, closer to the full model

Top-1 agreement with bf16 against file size: the files in this repo (green) sit above standard llama.cpp quantization (gray) at 2-bit, Q3 and Q4_K_M; at 3.9 GB, 87% against 70%

Most GGUF files come from one llama-quantize pass that rounds every weight to the target type; our 2-bit, Q3 and Q4_K_M files go further. In the 2-bit and Q3 files, we measured how much each weight tensor hurts the model when it is squeezed and gave the sensitive tensors more bits and the robust ones fewer. All three were then trained layer by layer to reproduce the full model (quantization-aware training) and distilled from the full bf16 model, used as the teacher, on English and Chinese rewriting data. For the 2-bit file, a standard build of the same size picks the same next token as bf16 70% of the time; these steps take it to 81%, 85% and then 87%. The result is still an ordinary GGUF that any recent llama.cpp runs; Q6_K and Q8_0 are standard builds, because at those sizes there is almost nothing left to recover.

Top-1 = how often the file's most likely next token is the same as bf16's ("Same top p" from llama.cpp llama-perplexity --kl-divergence), on about 30,700 held-out tokens per language; English and Chinese have equal token counts and are averaged. Every point is a measured file. Data: assets/quant-top1.csv; more detail on GitHub.

Files

The files here are byte-for-byte the same as the GGUF files in the main repository; download from either.

File Size Peak memory (Mac) ยน Top-1 vs. bf16 ยฒ Quality Use it for
humanizer-12b-Q8_0.gguf 12.7 GB 13.7 GB 98.4% Best 32 GB of memory or more. Recommended.
humanizer-12b-Q6_K.gguf 10.0 GB about 11.0 GB 97.7% No measurable loss 16 GB
humanizer-12b-Q4_K_M.gguf 7.6 GB 10.0 GB 95.2% Slight loss; quantization-aware trained About 14 GB, or when disk is tight
humanizer-12b-IQ3_XXS-QAT.gguf (the Q3 size) ยณ 5.6 GB 8.0 GB 93.1% Small loss; a few more fact slips in English 12 GB; proofread numbers and names
humanizer-12b-IQ2_XS-QAT.gguf (2-bit) ยณ 3.9 GB 6.2 GB 87.4% Lowest AI-detector score; a few more fact slips 8 GB; proofread numbers and names
humanizer-12b-bf16.gguf 23.8 GB about 24.8 GB 100% (reference) Reference Quantizing it yourself; reference runs

prompt_format.json (tiny) holds the instruction and separator verbatim; you need it unless you use the chat template (see Usage).

ยน llama-server (Metal) on an M5 Max with the app's settings (8,192-token context, one request at a time), while rewriting; Q6_K and bf16 are estimates. Leave room for the system (the app keeps about 4 GB free).

ยฒ English and Chinese averaged; per-language numbers and KL are below.

ยณ Mixed precision, quantization-aware trained and distilled, with a pruned vocabulary; see How the files were made. Both are named after the base type in their header, but neither is a plain build of that type. humanizer-12b-IQ3_XXS-QAT.gguf here is the same file as humanizer-12b-Q3-QAT.gguf in the main repository (same sha256); the app downloads that copy.

Quality on the task

Every file, measured the same way: top-1 and KL against bf16, a strict fact judge on the held-out evaluation set, and an AI detector on the same 60 English drafts.

File Top-1 vs. bf16, EN / ZH KL to bf16, EN / ZH Fact check (GLM): rewrites flagged ยท spots listed โด Flagged as AI, same 60 drafts โต
bf16 100% (reference) 0 (reference) EN 52 of 420 ยท 160 spots
ZH 50 of 200 ยท 216 spots
4 / 60
Q8_0 98.4% / 98.3% 0.0017 / 0.0016 EN 44 ยท 135 spots
ZH 54 ยท 236 spots
7 / 60
Q6_K 98.0% / 97.5% 0.0031 / 0.0035 EN 56 ยท 153 spots
ZH 56 ยท 255 spots
not measured
Q4_K_M (quantization-aware) 95.6% / 94.8% 0.0136 / 0.0146 EN 58 ยท 162 spots
ZH 51 ยท 194 spots
not measured
Q3 (IQ3_XXS-QAT) 93.6% / 92.6% 0.030 / 0.032 EN 64 ยท 210 spots
ZH 53 ยท 233 spots
4 / 60
2-bit (IQ2_XS-QAT) 87.7% / 87.1% 0.106 / 0.106 EN 70 ยท 265 spots
ZH 66 ยท 361 spots
0 / 60

Top-1 and KL: llama.cpp llama-perplexity --kl-divergence, 30 chunks of 2,048 tokens per language (about 30,700 scored tokens each) of held-out drafts and rewrites (no overlap with calibration or training data), reference = the bf16 GGUF, measured on A100 and A30 GPUs (the same file measures within about 3% in KL across GPU models). KL is how far a file's next-token probabilities drift from bf16, averaged over every token; lower is better. The Q3 and 2-bit files are measured against a bf16 with the same pruned vocabulary (pruning on its own: KL 0.0013 English / 0.0002 Chinese).

โด A strict LLM judge (GLM-5.3, one vote per rewrite) compared every rewrite with its draft: 420 English and 204 Chinese rewrites from the held-out evaluation set (2 per draft), run on each file with llama.cpp; 4 Chinese rewrites from bf16 could not be judged. A second pass re-read each flagged rewrite and listed every problem ("spots"). About 9 in 10 spots are one word, number or short phrase for every file. Compared draft by draft, Q8_0, Q6_K and Q4_K_M are within noise of each other and of bf16, and the quantization-aware Q4_K_M is within noise of a standard Q4_K_M (English 56 ยท 163 spots, Chinese 48 ยท 218). The 2-bit file is not: in English it was flagged on 70 rewrites against 44 for Q8_0 on the same drafts, a real difference (in Chinese, 66 vs. 54, within noise). Q3 sits in between: in English 64 rewrites flagged, against 52 for bf16 and 58 for Q4_K_M; in Chinese 53, on par with bf16. The judge is deliberately strict and some flags are harmless rewording, but with any file, read numbers, dates and names before you send.

โต Originality.ai, strictest setting, same 60 English drafts. The 2-bit file had 0 of 60 flagged, against 7 of 60 for Q8_0 and 4 of 60 for bf16: all 7 drafts flagged for Q8_0 passed with the 2-bit file, and none went the other way (paired p = 0.016). It looks like the small drift that quantization adds makes the wording a little less predictable, so it reads less machine-like. The Q3 file had 4 of 60 flagged, the same as bf16. Q6_K and Q4_K_M were not measured. Detectors change; this is one measurement on one date.

Every point on the chart, and what each step adds

The gray line is plain llama.cpp quantization for this model: a llama-quantize run of each type with the same importance matrix our 2-bit and Q3 files started from, full vocabulary, every other setting at its default. Nothing is interpolated.

Line File / type Size Top-1, EN Top-1, ZH Average (plotted) KL, EN / ZH
this repo 2-bit (IQ2_XS-QAT) 3.89 GB 87.72% 87.07% 87.40% 0.106 / 0.106
this repo Q3 (IQ3_XXS-QAT) 5.59 GB 93.63% 92.64% 93.13% 0.030 / 0.032
this repo Q4_K_M, quantization-aware trained 7.63 GB 95.62% 94.77% 95.20% 0.0136 / 0.0146
this repo Q6_K (standard build, our calibration) 10.03 GB 97.95% 97.48% 97.72% 0.0031 / 0.0035
this repo Q8_0 (standard build) 12.67 GB 98.45% 98.25% 98.35% 0.0017 / 0.0016
standard IQ2_XXS 3.57 GB 67.89% 45.84% 56.86% 0.823 / 2.268
standard IQ2_XS 3.89 GB 74.44% 66.21% 70.32% 0.471 / 0.760
standard IQ2_M 4.37 GB 81.55% 76.78% 79.16% 0.245 / 0.367
standard Q2_K 4.83 GB 82.03% 78.84% 80.44% 0.224 / 0.303
standard IQ3_XXS 4.85 GB 85.70% 82.77% 84.23% 0.136 / 0.185
standard IQ3_M 5.73 GB 90.94% 89.19% 90.07% 0.056 / 0.067
standard IQ4_XS 6.64 GB 93.79% 92.58% 93.19% 0.026 / 0.029
standard Q4_K_M 7.38 GB 94.32% 93.45% 93.89% 0.021 / 0.023
standard Q5_K_M 8.55 GB 95.98% 95.62% 95.80% 0.011 / 0.011
standard Q6_K 9.79 GB 97.87% 97.41% 97.64% 0.0033 / 0.0035
standard Q8_0 12.67 GB 98.50% 98.24% 98.37% 0.0015 / 0.0019

Also measured, not on the chart: standard IQ1_M (3.20 GB, 52.86% / 39.10%), smaller than any file here, and standard Q3_K_M (6.09 GB, 89.98% / 88.23%), beaten by the smaller IQ3_M.

  • 2-bit: 17 points above the standard IQ2_XS of the same size, and above the standard IQ3_XXS that is 1 GB larger. Plain 2-bit quantization hurts Chinese much more than English (66.2% vs. 74.4%); in this file the two languages come out the same.
  • Q3: above the standard IQ3_M, which is a little larger (93.1% vs. 90.1%), and level with the standard IQ4_XS, which is 1.0 GB larger.
  • Q4_K_M: a real but small gain over the standard Q4_K_M (95.2% vs. 93.9%); a standard Q5_K_M (8.5 GB, 95.8%) is still a little closer to bf16. (The standard Q4_K_M here is 7.4 GB because llama-quantize's default keeps the embeddings at 6-bit; with 8-bit embeddings like ours it is 7.6 GB: 94.46% / 93.46%.)

What each step adds, at the same file size (top-1, English and Chinese averaged):

Standard build 1. Bits by sensitivity 2. + layer-by-layer QAT 3. + distillation (released)
2-bit, 3.9 GB 70.3% (IQ2_XS) 80.6% 85.4% 87.4%
Q3, 5.6 GB 90.1% (IQ3_M, 5.7 GB) 90.8% 92.5% 93.1%

The Q4_K_M file went through steps 2 and 3 on the standard Q4_K_M type (no mixed precision): 94.5% / 93.5% โ†’ 95.6% / 94.8% (English / Chinese), same size and format. Step 1 also drops vocabulary rows that English and Chinese never use, which frees a little room for the weights.

How the files were made

  • Our own calibration. Q6_K, Q4_K_M, Q3 and the 2-bit file use an importance matrix (imatrix) computed on our own English and Chinese rewriting data, not on generic web text (the one for Q3 and 2-bit also mixes in some general text). Q8_0 needs no imatrix.
  • Embeddings and output layer stay at 8-bit in Q8_0, Q6_K and Q4_K_M (the model ties them, so this is one tensor).
  • Mixed precision by sensitivity (Q3 and 2-bit). Starting from an all-low-bit build, each tensor was upgraded on its own to a higher-bit type and the drop in KL to bf16 was measured; the bits then went where they buy the most, within a file-size budget. The 2-bit file is, by bytes, about half IQ2_XS and a quarter IQ3_XXS, with Q4_K (including the embeddings) and a little 3- and 6-bit where it matters most: about 2.7 bits per weight on average, in the same file size as a standard IQ2_XS. The Q3 file is about half Q4_K and a fifth Q6_K, with IQ3_XXS (17%) and IQ2_XS (11%) on the least sensitive tensors: about 3.9 bits per weight.
  • Quantization-aware training (Q4_K_M, Q3 and 2-bit). Layer by layer, the quantized weights are trained to reproduce the output of the same layer in the bf16 model. What is trained is exactly what is saved: every exported file was checked value by value against the trained weights.
  • Distillation (Q4_K_M, Q3 and 2-bit). The bf16 model is the teacher: the quantization scales and norms of the whole quantized model are trained to match its next-token probabilities on English and Chinese rewriting data.
  • Vocabulary pruning (Q3 and 2-bit). The files keep about 130,000 of the 262,144 tokens: the ones English and Chinese text actually uses, plus everything needed to spell any input. Any text still encodes and decodes exactly; rare symbols, emoji and other scripts just take a few more tokens.
  • Ordinary GGUF. No custom kernels and no new tensor types: any recent llama.cpp, and apps built on it, run the files.

Run it in the app

The humanizer app: draft on the left, rewrite on the right, new wording highlighted

humanizer also comes as a desktop app for macOS (Apple silicon) and Windows (x64) that runs these GGUF files locally. It is a small local web app: double-click it and it opens in your browser at http://127.0.0.1, runs llama.cpp in the background (Metal on Mac; CUDA, Vulkan or CPU on Windows), and works offline once the model is downloaded. Your text never leaves your computer. Paste a draft on the left, get the rewrite on the right, with new wording highlighted and replaced wording struck through. Since 0.3.2 it can also import the text of a .docx or PDF file, and any passage you select and mark with Create fact comes back word for word.

  • Download: Humanizer-0.3.2-macos-arm64.dmg or Humanizer-0.3.2-windows-x64-setup.exe (or the portable .zip) from app 0.3.2 on GitHub Releases (later versions: latest release). The app is not code-signed yet; the one-time first-launch fix is in the install guide.
  • Picking a size: on first launch the setup page shows your memory and lists the sizes with a short quality note each, marking one as Recommended: 32 GB or more โ†’ Q8_0 ("Best quality"), 16 GB โ†’ Q6_K ("Nearly identical to Q8_0"), 14 GB โ†’ Q4_K_M ("Slight loss"), 12 GB โ†’ Q3 ("Small loss; a few more fact slips in English"), less โ†’ 2-bit ("Lowest AI-detector score; a few more fact slips"). The app picks the largest size that leaves about 4 GB for the system and your browser. You can pick any of them and download it once (downloads can be paused and resumed; Hugging Face or the hf-mirror.com mirror). To switch later, open the โ‹ฏ menu โ†’ Change model size.
  • Which files the app offers: all five quantized files in the table, Q8_0, Q6_K, Q4_K_M, Q3 and 2-bit (app 0.3.0 and later; 0.2.0 had only the first three). The bf16 file is not in the app; use it with llama.cpp as below.

Usage

Text completion, not chat. The model was trained on one plain prompt: the instruction from prompt_format.json, a blank line, your draft, then \n\n### Rewritten:\n\n. No system prompt, no chat turns. Sampling: temperature 1.0, top-p 0.95, and nothing else (top-k 0, min-p 0, repetition penalty 1.0); stop on EOS only, no stop strings.

brew install llama.cpp            # or: winget install llama.cpp / a zip from github.com/ggml-org/llama.cpp/releases
pip install -U "huggingface_hub[cli]"
hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json --local-dir ./humanizer-model
#   16 GB: humanizer-12b-Q6_K.gguf ยท tight: humanizer-12b-Q4_K_M.gguf ยท tighter: humanizer-12b-IQ3_XXS-QAT.gguf ยท smallest: humanizer-12b-IQ2_XS-QAT.gguf
llama-server -m ./humanizer-model/humanizer-12b-Q8_0.gguf -c 8192 -np 1 -ngl 99 --host 127.0.0.1 --port 8080
import json, urllib.request

PF = json.load(open("humanizer-model/prompt_format.json", encoding="utf-8"))

def humanize(draft: str, url: str = "http://127.0.0.1:8080") -> str:
    body = {"prompt": PF["instr"] + "\n\n" + draft.strip() + PF["sep"],
            "temperature": 1.0, "top_p": 0.95, "top_k": 0, "min_p": 0, "repeat_penalty": 1.0,
            "n_predict": 2048}                    # no "stop": the model ends at EOS
    req = urllib.request.Request(url + "/completion", json.dumps(body).encode("utf-8"),
                                 {"Content-Type": "application/json"})
    with urllib.request.urlopen(req, timeout=900) as r:
        return json.load(r)["content"].strip()

print(humanize(open("draft.txt", encoding="utf-8").read()))

Chat template and defaults. Every file here carries a chat template that builds exactly this prompt from the last user message (system prompts and earlier turns are ignored), and stores the sampling settings above as defaults. So llama-server's /v1/chat/completions (with --jinja, the default in recent builds) and chat apps that use the file's template, such as LM Studio, work with one draft per message. Ollama ignores the GGUF's template; use the Modelfile in USAGE.md, section 7. Keep instruction + draft + rewrite within 8,192 tokens and split long documents at paragraph breaks, or use the hz command-line tool, which does it for you.

Wrong language, occasionally. On short, informal English drafts with technical jargon, the model occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and hz check the language and sample again automatically. If you call the model yourself: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.

Full guide for every runtime, a batch script, long documents, Chinese and troubleshooting: USAGE.md. sha256 checksums:

File Bytes sha256
humanizer-12b-bf16.gguf 23,832,049,568 47d79b44c3e15ea2540f4edb63556f2b7e456067d7dd252a42ff103a26a51be9
humanizer-12b-Q8_0.gguf 12,669,630,368 8d7a457b56de6530eaaf0151ccfa7550a4b20dab979737e259da9c63e960e0b0
humanizer-12b-Q6_K.gguf 10,029,799,584 c98f03bb9e71456181f99b0e1d3391e07ce1afc357f9db4c33d6d379b8dd9f0d
humanizer-12b-Q4_K_M.gguf 7,625,160,864 2229574dec5178629575ee4a153dfee7d9e924ab997d0ad9d2ac622b67e44834
humanizer-12b-IQ3_XXS-QAT.gguf 5,587,794,816 307bbfdf66fb22bf98aaa93fe7d54a713bf167e646073dc0cf10870b1525eb94
humanizer-12b-IQ2_XS-QAT.gguf 3,893,632,896 383e5ca8f1f48ab5f65013adbc1965fa70d1d1afa6c45d932c34b854e38edbb5

Links

License

Apache License 2.0. Fine-tuned from google/gemma-4-12B, which Google releases under Apache 2.0. This project is not affiliated with or endorsed by Google.

Downloads last month
3,822
GGUF
Model size
12B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jialinyyzz/humanizer-GGUF

Quantized
(7)
this model

Spaces using jialinyyzz/humanizer-GGUF 2