Humanizer 12B — MLX 8-bit

Unofficial MLX text-generation conversion of jialinyyzz/humanizer, not a new fine-tune or an official release by the original authors, Google, or Apple.

Original model: https://huggingface.co/jialinyyzz/humanizer
Original code: https://github.com/sgaofen/humanizer-local-model
Exact source snapshot: 7232532cb8466bd393544e97826e12e6b5e471ac

Build

Property This package
Source Published, already-trained BF16 12B checkpoint; no adapter merge
Conversion Official mlx_lm.convert.convert, Linux CPU
Quantization MLX affine, 8-bit, group size 64
Architecture handling Native MLX-LM gemma4_unified → text-language-model implementation
Saved weight-file size 11.78 GiB, measured from this build; not peak RAM
MLX / MLX-LM 0.32.3 / 0.32.0
Validation structural_checks_passed_cpu_smoke_skipped_not_benchmarked

MLX's native converter handles tensor names and excludes non-text components. It retains floating-point tensors where its model-specific quantization rules require them. This is a text-only package: the source multimodal configuration does not imply tested vision/audio support. No vocabulary pruning, retraining, QAT, or GGUF dequantization was performed.

Use on Apple silicon

Use native ARM64 Python 3.11 or newer and an OS supported by the installed MLX wheel. The requirements file pins the versions used by this converter; Apple Metal still needs local validation.

python3 -m venv .venv
source .venv/bin/activate
pip install "mlx==0.32.3" "mlx-lm==0.32.0" "transformers==5.14.1" "huggingface_hub==1.33.0"
hf download cowWhySo/humanizer-12b-mlx-8bit --local-dir ./humanizer-mlx
python ./humanizer-mlx/humanize_mlx.py --model ./humanizer-mlx --input draft.txt --output rewritten.txt

The included runner accepts a UTF-8 file or stdin, uses a raw completion, refuses to overwrite an existing output, and rejects an input/output token budget above 8,192 rather than silently truncating the draft. Split long documents at paragraph boundaries.

Python API

import json
from pathlib import Path
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

path = Path("humanizer-mlx")
pf = json.loads((path / "prompt_format.json").read_text(encoding="utf-8"))
model, tokenizer = load(str(path), tokenizer_config={"trust_remote_code": False}, trust_remote_code=False)
draft = Path("draft.txt").read_text(encoding="utf-8")
prompt = pf["instr"] + "\n\n" + draft.strip() + pf["sep"]
if len(tokenizer.encode(prompt)) + 2048 > 8192:
    raise ValueError("Split the document or lower the output-token allowance")
rewrite = generate(model, tokenizer, prompt=prompt, max_tokens=2048,
                   sampler=make_sampler(temp=1.0, top_p=0.95, top_k=0, min_p=0.0),
                   verbose=False)
print(rewrite.strip())

Do not apply a generic Gemma/chat template. The mlx_lm.generate CLI requires --ignore-chat-template; the Python code above passes the plain string directly. prompt_format.json is copied byte for byte from the pinned source. Its probe prompt for draft X must have SHA-256 prefix cc51d66b4c593fbe. English and Chinese drafts use the same instruction. Stop at EOS; do not add a ### stop string.

The output generation_config.json explicitly uses temperature 1.0, top-p 0.95, top-k 0, min-p 0.0, and repetition penalty 1.0. This corrects the source file's top-k 64 to the settings in the author's usage guide. The unmodified source configuration is retained in provenance/upstream_generation_config.json.

Validation and limitations

Safetensors byte spans/shard indexes, quantization metadata, strict MLX parameter loading, all floating-point parameters for finiteness, exact prompt bytes, and token IDs on a small probe set were checked. Full-checkpoint CPU generation diagnostics were disabled.

These are conversion and execution checks, not proof of BF16 logit parity, unchanged rewriting quality, fact preservation, latency, or peak memory. No Apple Metal test or full quality benchmark was performed by the Colab notebook. See validation.json for what actually ran. The CPU snippets contain only two generated tokens per draft, not complete evaluated rewrites.

Upstream reports SFT, DPO, and GRPO training; the public repository documents the recipe and data but says full training code will be released later. Its scripts/rebuild_reference_only.py rebuilds dataset source references, not model weights. Conversion does not require any training data.

The author's smaller GGUF releases include additional quantization-aware training and distillation. An MLX affine 8-bit build is not the same format or recipe as GGUF Q4_K_M, Q6_K, Q8_0, IQ3, or IQ2; their measured results do not transfer to this package.

The bundled runner does not implement the app's anti-copy resampling, language retries, Markdown/document-layout protection, or fact checks. Those are inference/application policies, not missing model weights. Review all output, particularly numbers, names, dates, units, quotations, and possible truncation. There is no guarantee of detector avoidance.

Provenance and license

The original LICENSE and NOTICE are included unchanged. The source release declares Apache-2.0; conversion does not remove its attribution or notice requirements.

BUILD_INFO.json describes intentional changes. PACKAGE_MANIFEST.json records local file hashes. provenance/source_lock.json records the immutable source commit, source file metadata, conversion settings, helper hashes, and library versions. The CPU environment freeze and source verification records are also included. Upload integrity is verified against the committed Hub file hashes before a newly created private repository is made public.

Primary references

Downloads last month
-
Safetensors
Model size
12B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for npario/humanizer-12b-mlx-8bit

Quantized
(7)
this model