Instructions to use npario/humanizer-12b-mlx-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use npario/humanizer-12b-mlx-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("npario/humanizer-12b-mlx-8bit") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use npario/humanizer-12b-mlx-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "npario/humanizer-12b-mlx-8bit" --prompt "Once upon a time"
- Atomic Chat
Humanizer 12B — MLX 8-bit
Unofficial MLX text-generation conversion of jialinyyzz/humanizer, not a new fine-tune or an official release by the original authors, Google, or Apple.
Original model: https://huggingface.co/jialinyyzz/humanizer
Original code: https://github.com/sgaofen/humanizer-local-model
Exact source snapshot: 7232532cb8466bd393544e97826e12e6b5e471ac
Build
| Property | This package |
|---|---|
| Source | Published, already-trained BF16 12B checkpoint; no adapter merge |
| Conversion | Official mlx_lm.convert.convert, Linux CPU |
| Quantization | MLX affine, 8-bit, group size 64 |
| Architecture handling | Native MLX-LM gemma4_unified → text-language-model implementation |
| Saved weight-file size | 11.78 GiB, measured from this build; not peak RAM |
| MLX / MLX-LM | 0.32.3 / 0.32.0 |
| Validation | structural_checks_passed_cpu_smoke_skipped_not_benchmarked |
MLX's native converter handles tensor names and excludes non-text components. It retains floating-point tensors where its model-specific quantization rules require them. This is a text-only package: the source multimodal configuration does not imply tested vision/audio support. No vocabulary pruning, retraining, QAT, or GGUF dequantization was performed.
Use on Apple silicon
Use native ARM64 Python 3.11 or newer and an OS supported by the installed MLX wheel. The requirements file pins the versions used by this converter; Apple Metal still needs local validation.
python3 -m venv .venv
source .venv/bin/activate
pip install "mlx==0.32.3" "mlx-lm==0.32.0" "transformers==5.14.1" "huggingface_hub==1.33.0"
hf download cowWhySo/humanizer-12b-mlx-8bit --local-dir ./humanizer-mlx
python ./humanizer-mlx/humanize_mlx.py --model ./humanizer-mlx --input draft.txt --output rewritten.txt
The included runner accepts a UTF-8 file or stdin, uses a raw completion, refuses to overwrite an existing output, and rejects an input/output token budget above 8,192 rather than silently truncating the draft. Split long documents at paragraph boundaries.
Python API
import json
from pathlib import Path
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
path = Path("humanizer-mlx")
pf = json.loads((path / "prompt_format.json").read_text(encoding="utf-8"))
model, tokenizer = load(str(path), tokenizer_config={"trust_remote_code": False}, trust_remote_code=False)
draft = Path("draft.txt").read_text(encoding="utf-8")
prompt = pf["instr"] + "\n\n" + draft.strip() + pf["sep"]
if len(tokenizer.encode(prompt)) + 2048 > 8192:
raise ValueError("Split the document or lower the output-token allowance")
rewrite = generate(model, tokenizer, prompt=prompt, max_tokens=2048,
sampler=make_sampler(temp=1.0, top_p=0.95, top_k=0, min_p=0.0),
verbose=False)
print(rewrite.strip())
Do not apply a generic Gemma/chat template. The mlx_lm.generate CLI requires
--ignore-chat-template; the Python code above passes the plain string directly.
prompt_format.json is copied byte for byte from the pinned source. Its probe prompt for draft
X must have SHA-256 prefix cc51d66b4c593fbe. English and Chinese drafts use the same instruction.
Stop at EOS; do not add a ### stop string.
The output generation_config.json explicitly uses temperature 1.0, top-p 0.95, top-k 0,
min-p 0.0, and repetition penalty 1.0. This corrects the source file's top-k 64 to the
settings in the author's usage guide. The unmodified source configuration is retained in
provenance/upstream_generation_config.json.
Validation and limitations
Safetensors byte spans/shard indexes, quantization metadata, strict MLX parameter loading, all floating-point parameters for finiteness, exact prompt bytes, and token IDs on a small probe set were checked. Full-checkpoint CPU generation diagnostics were disabled.
These are conversion and execution checks, not proof of BF16 logit parity, unchanged
rewriting quality, fact preservation, latency, or peak memory. No Apple Metal test or full
quality benchmark was performed by the Colab notebook. See validation.json for what actually ran.
The CPU snippets contain only two generated tokens per draft, not complete evaluated rewrites.
Upstream reports SFT, DPO, and GRPO training; the public repository documents the recipe and
data but says full training code will be released later. Its scripts/rebuild_reference_only.py
rebuilds dataset source references, not model weights. Conversion does not require any training data.
The author's smaller GGUF releases include additional quantization-aware training and distillation. An MLX affine 8-bit build is not the same format or recipe as GGUF Q4_K_M, Q6_K, Q8_0, IQ3, or IQ2; their measured results do not transfer to this package.
The bundled runner does not implement the app's anti-copy resampling, language retries, Markdown/document-layout protection, or fact checks. Those are inference/application policies, not missing model weights. Review all output, particularly numbers, names, dates, units, quotations, and possible truncation. There is no guarantee of detector avoidance.
Provenance and license
The original LICENSE and NOTICE are included unchanged. The source release declares
Apache-2.0; conversion does not remove its attribution or notice requirements.
BUILD_INFO.json describes intentional changes. PACKAGE_MANIFEST.json records local file
hashes. provenance/source_lock.json records the immutable source commit, source file metadata,
conversion settings, helper hashes, and library versions. The CPU environment freeze and source
verification records are also included. Upload integrity is verified against the committed Hub
file hashes before a newly created private repository is made public.
Primary references
- Downloads last month
- -
8-bit