Llama-3.2-1B-RYS-10-13-GGUF

A layer-duplication ("RYS" — Repeat Your Self, David Ng) variant of meta-llama/Llama-3.2-1B-Instruct: transformer layers 10–12 are duplicated, expanding the stack from 16 to 19 layers. No training, no merging, no weight changes — purely structural duplication. GGUF (imatrix Q-quants).

⚠️ Evaluation status — please read (updated 2026-06)

The large reasoning gain originally reported for this model was a measurement artifact, not a real capability gain. This card is being corrected to say so plainly.

The original card reported reasoning 0.00% → 64.71%. That 0% baseline came from a degraded inference setup, not from the model. On a current llama.cpp build the unmodified Llama-3.2-1B-Instruct already scores about 52.94% on the same reasoning probe, and this (10,13) duplication adds ~0 over that baseline (and slightly lowers the EQ probe). The headline "+64.71" was the old stack's broken floor rising back to normal — not something the duplication unlocked.

Why: the scores come from a lightweight search probe (16 math / 16 EQ / 17 reasoning questions, greedy-decoded) used to locate productive layer blocks — not a validated benchmark. Reasoning moves in steps of 1/17 ≈ 5.9%, so the deltas are coarse, and a degraded baseline can manufacture a huge apparent gain.

Independent benchmark (lm-eval-harness, GSM8K 5-shot, N=100): base 34% strict-match (38% flexible); RYS (10,13) 23% strict (27% flexible) — same 100 problems. So on a real benchmark the duplication does not improve reasoning and if anything lowers it. This both confirms the base is far from "0%" and refutes the "+64.71pp" claim directionally. (N=100; a larger paired run would tighten the magnitude, but the direction is clear.)

Bottom line: treat this as a normal Llama-3.2-1B-Instruct with layers 10–12 duplicated. Published for transparency and reproducibility of the RYS sweep — not as an improved reasoner.

Original sweep numbers (search probe — kept for the record)

probe reported baseline reported (10,13) re-test note
Reasoning (17 q) 0.00% 64.71% baseline was a degraded-stack artifact; correct-stack baseline ≈ 52.94%, (10,13) Δ ≈ 0. Real GSM8K (N=100): base 34%, (10,13) 23% — duplication lowers it
EQ (16 q) 27.11 90.12 baseline also stack-dependent; not a validated EQ benchmark
Math (16 q) 0.536 0.711 search-probe score, unconfirmed

Run it

llama-server -m Llama-3.2-1B-RYS-10-13-Q4_K_M.gguf -ngl 99

Method · data · attribution

  • Method: layer duplication — Repeat Your Self (David Ng); toolkit llm-circuit-finder (alainnothere).
  • Raw sweep data: rys-sovereign-collection-v2.
  • Built by John Broadway with Claude. The method and the raw data are real and reproducible; the interpretation of the original probe deltas as capability is what this update corrects.

License

Llama 3.2 Community License (inherits from the base model).

Downloads last month
14
GGUF
Model size
1B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for john-broadway/Llama-3.2-1B-RYS-10-13-GGUF

Quantized
(423)
this model

Collection including john-broadway/Llama-3.2-1B-RYS-10-13-GGUF