Instructions to use jialinyyzz/humanizer-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jialinyyzz/humanizer-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jialinyyzz/humanizer-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jialinyyzz/humanizer-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jialinyyzz/humanizer-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jialinyyzz/humanizer-GGUF:Q4_K_M
Use Docker
docker model run hf.co/jialinyyzz/humanizer-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use jialinyyzz/humanizer-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jialinyyzz/humanizer-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jialinyyzz/humanizer-GGUF:Q4_K_M
- Ollama
How to use jialinyyzz/humanizer-GGUF with Ollama:
ollama run hf.co/jialinyyzz/humanizer-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use jialinyyzz/humanizer-GGUF with Docker Model Runner:
docker model run hf.co/jialinyyzz/humanizer-GGUF:Q4_K_M
- Lemonade
How to use jialinyyzz/humanizer-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jialinyyzz/humanizer-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.humanizer-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
humanizer 12B: GGUF quantizations
GGUF files of jialinyyzz/humanizer (v2), a 12B model that rewrites AI-written drafts in English and Chinese so they read like a person wrote them, trained to keep every number, date, name and quote. The 2-bit, Q3 and Q4_K_M files are quantization-aware trained and distilled from bf16, so they stay closer to the full model than standard quantizations of the same size.
Quick start: the desktop app (Mac, Windows) downloads these files for you; or hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json and see Usage.
Results
95% judged human by Originality.ai at its strictest setting (210 English drafts, bf16 weights, 2026-10-02): 11 of 210 rewrites were flagged as AI, against 26 for the previous release, and no AI detector was used anywhere in training. On the same 60 drafts, the 2-bit file here had 0 of 60 flagged (Q8_0 7, bf16 4); per-file results are under Quality on the task. 376 of 420 English rewrites came back with no factual problem from a strict LLM judge, measured on the Q8_0 file here (210 drafts, two samples each); where it did find one, more than 9 in 10 fixes are a single word or phrase. Full results: main model card.
Smaller files, closer to the full model
Most GGUF files come from one llama-quantize pass that rounds every weight to the target type; our 2-bit, Q3 and Q4_K_M files go further. In the 2-bit and Q3 files, we measured how much each weight tensor hurts the model when it is squeezed and gave the sensitive tensors more bits and the robust ones fewer. All three were then trained layer by layer to reproduce the full model (quantization-aware training) and distilled from the full bf16 model, used as the teacher, on English and Chinese rewriting data. For the 2-bit file, a standard build of the same size picks the same next token as bf16 70% of the time; these steps take it to 81%, 85% and then 87%. The result is still an ordinary GGUF that any recent llama.cpp runs; Q6_K and Q8_0 are standard builds, because at those sizes there is almost nothing left to recover.
Top-1 = how often the file's most likely next token is the same as bf16's ("Same top p" from llama.cpp llama-perplexity --kl-divergence), on about 30,700 held-out tokens per language; English and Chinese have equal token counts and are averaged. Every point is a measured file. Data: assets/quant-top1.csv; more detail on GitHub.
Files
The files here are byte-for-byte the same as the GGUF files in the main repository; download from either.
| File | Size | Peak memory (Mac) ยน | Top-1 vs. bf16 ยฒ | Quality | Use it for |
|---|---|---|---|---|---|
humanizer-12b-Q8_0.gguf |
12.7 GB | 13.7 GB | 98.4% | Best | 32 GB of memory or more. Recommended. |
humanizer-12b-Q6_K.gguf |
10.0 GB | about 11.0 GB | 97.7% | No measurable loss | 16 GB |
humanizer-12b-Q4_K_M.gguf |
7.6 GB | 10.0 GB | 95.2% | Slight loss; quantization-aware trained | About 14 GB, or when disk is tight |
humanizer-12b-IQ3_XXS-QAT.gguf (the Q3 size) ยณ |
5.6 GB | 8.0 GB | 93.1% | Small loss; a few more fact slips in English | 12 GB; proofread numbers and names |
humanizer-12b-IQ2_XS-QAT.gguf (2-bit) ยณ |
3.9 GB | 6.2 GB | 87.4% | Lowest AI-detector score; a few more fact slips | 8 GB; proofread numbers and names |
humanizer-12b-bf16.gguf |
23.8 GB | about 24.8 GB | 100% (reference) | Reference | Quantizing it yourself; reference runs |
prompt_format.json (tiny) holds the instruction and separator verbatim; you need it unless you use the chat template (see Usage).
ยน llama-server (Metal) on an M5 Max with the app's settings (8,192-token context, one request at a time), while rewriting; Q6_K and bf16 are estimates. Leave room for the system (the app keeps about 4 GB free).
ยฒ English and Chinese averaged; per-language numbers and KL are below.
ยณ Mixed precision, quantization-aware trained and distilled, with a pruned vocabulary; see How the files were made. Both are named after the base type in their header, but neither is a plain build of that type. humanizer-12b-IQ3_XXS-QAT.gguf here is the same file as humanizer-12b-Q3-QAT.gguf in the main repository (same sha256); the app downloads that copy.
Quality on the task
Every file, measured the same way: top-1 and KL against bf16, a strict fact judge on the held-out evaluation set, and an AI detector on the same 60 English drafts.
| File | Top-1 vs. bf16, EN / ZH | KL to bf16, EN / ZH | Fact check (GLM): rewrites flagged ยท spots listed โด | Flagged as AI, same 60 drafts โต |
|---|---|---|---|---|
| bf16 | 100% (reference) | 0 (reference) | EN 52 of 420 ยท 160 spots ZH 50 of 200 ยท 216 spots |
4 / 60 |
| Q8_0 | 98.4% / 98.3% | 0.0017 / 0.0016 | EN 44 ยท 135 spots ZH 54 ยท 236 spots |
7 / 60 |
| Q6_K | 98.0% / 97.5% | 0.0031 / 0.0035 | EN 56 ยท 153 spots ZH 56 ยท 255 spots |
not measured |
| Q4_K_M (quantization-aware) | 95.6% / 94.8% | 0.0136 / 0.0146 | EN 58 ยท 162 spots ZH 51 ยท 194 spots |
not measured |
Q3 (IQ3_XXS-QAT) |
93.6% / 92.6% | 0.030 / 0.032 | EN 64 ยท 210 spots ZH 53 ยท 233 spots |
4 / 60 |
2-bit (IQ2_XS-QAT) |
87.7% / 87.1% | 0.106 / 0.106 | EN 70 ยท 265 spots ZH 66 ยท 361 spots |
0 / 60 |
Top-1 and KL: llama.cpp llama-perplexity --kl-divergence, 30 chunks of 2,048 tokens per language (about 30,700 scored tokens each) of held-out drafts and rewrites (no overlap with calibration or training data), reference = the bf16 GGUF, measured on A100 and A30 GPUs (the same file measures within about 3% in KL across GPU models). KL is how far a file's next-token probabilities drift from bf16, averaged over every token; lower is better. The Q3 and 2-bit files are measured against a bf16 with the same pruned vocabulary (pruning on its own: KL 0.0013 English / 0.0002 Chinese).
โด A strict LLM judge (GLM-5.3, one vote per rewrite) compared every rewrite with its draft: 420 English and 204 Chinese rewrites from the held-out evaluation set (2 per draft), run on each file with llama.cpp; 4 Chinese rewrites from bf16 could not be judged. A second pass re-read each flagged rewrite and listed every problem ("spots"). About 9 in 10 spots are one word, number or short phrase for every file. Compared draft by draft, Q8_0, Q6_K and Q4_K_M are within noise of each other and of bf16, and the quantization-aware Q4_K_M is within noise of a standard Q4_K_M (English 56 ยท 163 spots, Chinese 48 ยท 218). The 2-bit file is not: in English it was flagged on 70 rewrites against 44 for Q8_0 on the same drafts, a real difference (in Chinese, 66 vs. 54, within noise). Q3 sits in between: in English 64 rewrites flagged, against 52 for bf16 and 58 for Q4_K_M; in Chinese 53, on par with bf16. The judge is deliberately strict and some flags are harmless rewording, but with any file, read numbers, dates and names before you send.
โต Originality.ai, strictest setting, same 60 English drafts. The 2-bit file had 0 of 60 flagged, against 7 of 60 for Q8_0 and 4 of 60 for bf16: all 7 drafts flagged for Q8_0 passed with the 2-bit file, and none went the other way (paired p = 0.016). It looks like the small drift that quantization adds makes the wording a little less predictable, so it reads less machine-like. The Q3 file had 4 of 60 flagged, the same as bf16. Q6_K and Q4_K_M were not measured. Detectors change; this is one measurement on one date.
Every point on the chart, and what each step adds
The gray line is plain llama.cpp quantization for this model: a llama-quantize run of each type with the same importance matrix our 2-bit and Q3 files started from, full vocabulary, every other setting at its default. Nothing is interpolated.
| Line | File / type | Size | Top-1, EN | Top-1, ZH | Average (plotted) | KL, EN / ZH |
|---|---|---|---|---|---|---|
| this repo | 2-bit (IQ2_XS-QAT) |
3.89 GB | 87.72% | 87.07% | 87.40% | 0.106 / 0.106 |
| this repo | Q3 (IQ3_XXS-QAT) |
5.59 GB | 93.63% | 92.64% | 93.13% | 0.030 / 0.032 |
| this repo | Q4_K_M, quantization-aware trained | 7.63 GB | 95.62% | 94.77% | 95.20% | 0.0136 / 0.0146 |
| this repo | Q6_K (standard build, our calibration) | 10.03 GB | 97.95% | 97.48% | 97.72% | 0.0031 / 0.0035 |
| this repo | Q8_0 (standard build) | 12.67 GB | 98.45% | 98.25% | 98.35% | 0.0017 / 0.0016 |
| standard | IQ2_XXS | 3.57 GB | 67.89% | 45.84% | 56.86% | 0.823 / 2.268 |
| standard | IQ2_XS | 3.89 GB | 74.44% | 66.21% | 70.32% | 0.471 / 0.760 |
| standard | IQ2_M | 4.37 GB | 81.55% | 76.78% | 79.16% | 0.245 / 0.367 |
| standard | Q2_K | 4.83 GB | 82.03% | 78.84% | 80.44% | 0.224 / 0.303 |
| standard | IQ3_XXS | 4.85 GB | 85.70% | 82.77% | 84.23% | 0.136 / 0.185 |
| standard | IQ3_M | 5.73 GB | 90.94% | 89.19% | 90.07% | 0.056 / 0.067 |
| standard | IQ4_XS | 6.64 GB | 93.79% | 92.58% | 93.19% | 0.026 / 0.029 |
| standard | Q4_K_M | 7.38 GB | 94.32% | 93.45% | 93.89% | 0.021 / 0.023 |
| standard | Q5_K_M | 8.55 GB | 95.98% | 95.62% | 95.80% | 0.011 / 0.011 |
| standard | Q6_K | 9.79 GB | 97.87% | 97.41% | 97.64% | 0.0033 / 0.0035 |
| standard | Q8_0 | 12.67 GB | 98.50% | 98.24% | 98.37% | 0.0015 / 0.0019 |
Also measured, not on the chart: standard IQ1_M (3.20 GB, 52.86% / 39.10%), smaller than any file here, and standard Q3_K_M (6.09 GB, 89.98% / 88.23%), beaten by the smaller IQ3_M.
- 2-bit: 17 points above the standard IQ2_XS of the same size, and above the standard IQ3_XXS that is 1 GB larger. Plain 2-bit quantization hurts Chinese much more than English (66.2% vs. 74.4%); in this file the two languages come out the same.
- Q3: above the standard IQ3_M, which is a little larger (93.1% vs. 90.1%), and level with the standard IQ4_XS, which is 1.0 GB larger.
- Q4_K_M: a real but small gain over the standard Q4_K_M (95.2% vs. 93.9%); a standard Q5_K_M (8.5 GB, 95.8%) is still a little closer to bf16. (The standard Q4_K_M here is 7.4 GB because llama-quantize's default keeps the embeddings at 6-bit; with 8-bit embeddings like ours it is 7.6 GB: 94.46% / 93.46%.)
What each step adds, at the same file size (top-1, English and Chinese averaged):
| Standard build | 1. Bits by sensitivity | 2. + layer-by-layer QAT | 3. + distillation (released) | |
|---|---|---|---|---|
| 2-bit, 3.9 GB | 70.3% (IQ2_XS) | 80.6% | 85.4% | 87.4% |
| Q3, 5.6 GB | 90.1% (IQ3_M, 5.7 GB) | 90.8% | 92.5% | 93.1% |
The Q4_K_M file went through steps 2 and 3 on the standard Q4_K_M type (no mixed precision): 94.5% / 93.5% โ 95.6% / 94.8% (English / Chinese), same size and format. Step 1 also drops vocabulary rows that English and Chinese never use, which frees a little room for the weights.
How the files were made
- Our own calibration. Q6_K, Q4_K_M, Q3 and the 2-bit file use an importance matrix (imatrix) computed on our own English and Chinese rewriting data, not on generic web text (the one for Q3 and 2-bit also mixes in some general text). Q8_0 needs no imatrix.
- Embeddings and output layer stay at 8-bit in Q8_0, Q6_K and Q4_K_M (the model ties them, so this is one tensor).
- Mixed precision by sensitivity (Q3 and 2-bit). Starting from an all-low-bit build, each tensor was upgraded on its own to a higher-bit type and the drop in KL to bf16 was measured; the bits then went where they buy the most, within a file-size budget. The 2-bit file is, by bytes, about half IQ2_XS and a quarter IQ3_XXS, with Q4_K (including the embeddings) and a little 3- and 6-bit where it matters most: about 2.7 bits per weight on average, in the same file size as a standard IQ2_XS. The Q3 file is about half Q4_K and a fifth Q6_K, with IQ3_XXS (17%) and IQ2_XS (11%) on the least sensitive tensors: about 3.9 bits per weight.
- Quantization-aware training (Q4_K_M, Q3 and 2-bit). Layer by layer, the quantized weights are trained to reproduce the output of the same layer in the bf16 model. What is trained is exactly what is saved: every exported file was checked value by value against the trained weights.
- Distillation (Q4_K_M, Q3 and 2-bit). The bf16 model is the teacher: the quantization scales and norms of the whole quantized model are trained to match its next-token probabilities on English and Chinese rewriting data.
- Vocabulary pruning (Q3 and 2-bit). The files keep about 130,000 of the 262,144 tokens: the ones English and Chinese text actually uses, plus everything needed to spell any input. Any text still encodes and decodes exactly; rare symbols, emoji and other scripts just take a few more tokens.
- Ordinary GGUF. No custom kernels and no new tensor types: any recent llama.cpp, and apps built on it, run the files.
Run it in the app
humanizer also comes as a desktop app for macOS (Apple silicon) and Windows (x64) that runs these GGUF files locally. It is a small local web app: double-click it and it opens in your browser at http://127.0.0.1, runs llama.cpp in the background (Metal on Mac; CUDA, Vulkan or CPU on Windows), and works offline once the model is downloaded. Your text never leaves your computer. Paste a draft on the left, get the rewrite on the right, with new wording highlighted and replaced wording struck through. Since 0.3.2 it can also import the text of a .docx or PDF file, and any passage you select and mark with Create fact comes back word for word.
- Download:
Humanizer-0.3.2-macos-arm64.dmgorHumanizer-0.3.2-windows-x64-setup.exe(or the portable.zip) from app 0.3.2 on GitHub Releases (later versions: latest release). The app is not code-signed yet; the one-time first-launch fix is in the install guide. - Picking a size: on first launch the setup page shows your memory and lists the sizes with a short quality note each, marking one as Recommended: 32 GB or more โ Q8_0 ("Best quality"), 16 GB โ Q6_K ("Nearly identical to Q8_0"), 14 GB โ Q4_K_M ("Slight loss"), 12 GB โ Q3 ("Small loss; a few more fact slips in English"), less โ 2-bit ("Lowest AI-detector score; a few more fact slips"). The app picks the largest size that leaves about 4 GB for the system and your browser. You can pick any of them and download it once (downloads can be paused and resumed; Hugging Face or the hf-mirror.com mirror). To switch later, open the โฏ menu โ Change model size.
- Which files the app offers: all five quantized files in the table, Q8_0, Q6_K, Q4_K_M, Q3 and 2-bit (app 0.3.0 and later; 0.2.0 had only the first three). The bf16 file is not in the app; use it with llama.cpp as below.
Usage
Text completion, not chat. The model was trained on one plain prompt: the instruction from prompt_format.json, a blank line, your draft, then \n\n### Rewritten:\n\n. No system prompt, no chat turns. Sampling: temperature 1.0, top-p 0.95, and nothing else (top-k 0, min-p 0, repetition penalty 1.0); stop on EOS only, no stop strings.
brew install llama.cpp # or: winget install llama.cpp / a zip from github.com/ggml-org/llama.cpp/releases
pip install -U "huggingface_hub[cli]"
hf download jialinyyzz/humanizer-GGUF humanizer-12b-Q8_0.gguf prompt_format.json --local-dir ./humanizer-model
# 16 GB: humanizer-12b-Q6_K.gguf ยท tight: humanizer-12b-Q4_K_M.gguf ยท tighter: humanizer-12b-IQ3_XXS-QAT.gguf ยท smallest: humanizer-12b-IQ2_XS-QAT.gguf
llama-server -m ./humanizer-model/humanizer-12b-Q8_0.gguf -c 8192 -np 1 -ngl 99 --host 127.0.0.1 --port 8080
import json, urllib.request
PF = json.load(open("humanizer-model/prompt_format.json", encoding="utf-8"))
def humanize(draft: str, url: str = "http://127.0.0.1:8080") -> str:
body = {"prompt": PF["instr"] + "\n\n" + draft.strip() + PF["sep"],
"temperature": 1.0, "top_p": 0.95, "top_k": 0, "min_p": 0, "repeat_penalty": 1.0,
"n_predict": 2048} # no "stop": the model ends at EOS
req = urllib.request.Request(url + "/completion", json.dumps(body).encode("utf-8"),
{"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=900) as r:
return json.load(r)["content"].strip()
print(humanize(open("draft.txt", encoding="utf-8").read()))
Chat template and defaults. Every file here carries a chat template that builds exactly this prompt from the last user message (system prompts and earlier turns are ignored), and stores the sampling settings above as defaults. So llama-server's /v1/chat/completions (with --jinja, the default in recent builds) and chat apps that use the file's template, such as LM Studio, work with one draft per message. Ollama ignores the GGUF's template; use the Modelfile in USAGE.md, section 7. Keep instruction + draft + rewrite within 8,192 tokens and split long documents at paragraph breaks, or use the hz command-line tool, which does it for you.
Wrong language, occasionally. On short, informal English drafts with technical jargon, the model occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and hz check the language and sample again automatically. If you call the model yourself: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.
Full guide for every runtime, a batch script, long documents, Chinese and troubleshooting: USAGE.md. sha256 checksums:
| File | Bytes | sha256 |
|---|---|---|
humanizer-12b-bf16.gguf |
23,832,049,568 | 47d79b44c3e15ea2540f4edb63556f2b7e456067d7dd252a42ff103a26a51be9 |
humanizer-12b-Q8_0.gguf |
12,669,630,368 | 8d7a457b56de6530eaaf0151ccfa7550a4b20dab979737e259da9c63e960e0b0 |
humanizer-12b-Q6_K.gguf |
10,029,799,584 | c98f03bb9e71456181f99b0e1d3391e07ce1afc357f9db4c33d6d379b8dd9f0d |
humanizer-12b-Q4_K_M.gguf |
7,625,160,864 | 2229574dec5178629575ee4a153dfee7d9e924ab997d0ad9d2ac622b67e44834 |
humanizer-12b-IQ3_XXS-QAT.gguf |
5,587,794,816 | 307bbfdf66fb22bf98aaa93fe7d54a713bf167e646073dc0cf10870b1525eb94 |
humanizer-12b-IQ2_XS-QAT.gguf |
3,893,632,896 | 383e5ca8f1f48ab5f65013adbc1965fa70d1d1afa6c45d932c34b854e38edbb5 |
Links
- Main model (bf16 safetensors, training, full results): jialinyyzz/humanizer
- Code, app, evaluation data: github.com/sgaofen/humanizer-local-model
- Desktop app: Releases
License
Apache License 2.0. Fine-tuned from google/gemma-4-12B, which Google releases under Apache 2.0. This project is not affiliated with or endorsed by Google.
- Downloads last month
- 3,822
2-bit
3-bit
4-bit
6-bit
8-bit
16-bit