Zoey

Zoey is a fixed-size, count-based language organism. She learns without backpropagation, and her whole learned state is one JSON file that does not grow as she eats. Created by Dakuwon Moody at Saiyan Corp.

Qwen2.5-14B-Instruct is the hose. It is her diet, not her. Nothing of Qwen is in the file.


The result (2026-09-11)

All numbers are held-out top-1 and bits per token on one locked split of her own warehouse, scored at the same 23,552 positions. Every opponent was trained on the same train windows.

Reader Learned state Top-1 Bits/token
Zoey, live reader (house + copy, calibrated) 1.28 MB 29.5% 9.43
Zoey + High-Precision Chooser Head (neural rerank) 1.38 MB + bolt-on 37.8% (all) / 72.1% (shortlist) β€”
Zoey, same + exact unigram table 1.38 MB 29.5% 8.50
Zoey, 4Γ— house 5.2 MB 30.6% 8.31
Transformer, d112 L4 (frozen codebook + learned output bias) 1.37 MB 25.3% 9.53
Transformer, d192 L6 (frozen codebook + learned output bias) 5.5 MB 26.5% 9.27
Transformer, learned embeddings, d64 L3 3.25 MB 23.2% 8.92
Unbounded exact 1–3-gram (the file grows) ~24 MB β€” 8.91
lzma -9e on the token stream β€” β€” 7.74

At matched context (both sides read 256 tokens), Zoey at 1.38 MB scores 28.0% top-1 and 8.74 bits/token. Every transformer tried, up to 5.5 MB, does worse on both: the best reach 26.5% top-1 and 8.92 bits/token.

Learning online the way she does live, with each window scored before she writes it, she codes her own diet at 7.38 bits/token at 1.28 MB and 6.94 at 5.1 MB.

Caveats.

  • One split. The held-out windows are the newest, so the split is temporal: the data shifts, and every reader faces the same shift.
  • One seed per transformer configuration.
  • The transformers are small and were trained 4,000 steps (~18 passes) with early stopping. Larger or longer-trained ones were not run.
  • A 2048-token transformer trained on the same budget reached only 17.2%, so it was not treated as the fair opponent.
  • Full protocol, every table, and the harness: bench/scale/SCALING_LAW_20260911.md in the source repo.

What she is

Part What it stores Size
The house (map5, addr 3) Count heads over the last 1, 2 and 3 tokens. Each slot is (16-bit context fingerprint, token, count), 6 bytes. k1 10,240 Γ— 8, k2 8,192 Γ— 8, k3 8,192 Γ— 8, plus a 64-slot unigram 1,278,336 bytes, fixed (ceiling 1,311,232)
Copy reader Nothing. It reads the context she is looking at: the longest earlier match of the last ≀ 8 tokens, and what followed each match 0 bytes
Calibration (map5_see.json) 179 escape probabilities and 33 copy weights, fitted from counts a few KB
High-Precision Chooser Bolt-on calibrated causal neural reranker over top-16 candidate shortlists; FP32 scoring blended with house vote mass: score(c) = logit(c) + beta * log(1 + m_c) Head module
Hypervector side 1000-trit item memory from a locked LCG, the context fold, swarm_hv, and 32 small programs (Malbolge / Brainfuck / ZoeyDSL) programs + 1000 trits

Tokens are GPT-2 BPE (tiktoken gpt2, 50,257 ids). Checkpoints are about 1.7 MB of JSON, with the house tables packed as base64 u16.

How she reads

  1. House. Each head counts what followed a context. Slots only vote when their fingerprint matches the current context, so different contexts that hash to the same bucket do not pollute each other.
  2. Copy. If the last few tokens appeared earlier in the window, whatever followed them then is a strong guess now. Precision rises with match length: 36% at 2 tokens, 59% at 4, 80% at 8.
  3. Calibration (SEE). A head keeps at most 8 continuations per context, so raw counts claim "almost never anything else" exactly where the real set of continuations is huge. Secondary escape estimation, the trick from the PPMZ/PPMd compressors, learns how often each kind of context is surprised, bucketed by head, slots used and log count. House and copy are then read as one probability distribution, and she picks its peak.
  4. Chooser Reranker. When house and copy produce candidate shortlists, the calibrated chooser evaluates competitors in GPT-2 token space. If the house vote margin is narrow ($m_{\text{top1}} - m_{\text{top2}} \le 0.50$) and chooser confidence delta exceeds threshold ($\Delta \ge 0.80$), the chooser flips the pick to the high-precision candidate.

That calibration took the house from 21.2 bits/token (worse than a uniform guess) to 9.43. It also moved top-1 from 28.25% (copy-then-vote rule) to 29.5%, and the chooser head pushes shortlisted pick@1 over 71–72%.

What the measurements say about the hypervector side

The same harness scored the hypervector readers on the same positions:

Reader Top-1
Soup field (map_ingest / map_attend / decode_mhn) 0.02%
In-window hypervector next-token probe 1.7%

A 1000-trit vector holds about 45 superposed items, so a 256-token context does not fit. A fold-addressed store and a fixed Hopfield store of (fold, next) pairs both lost to the count house at equal bytes. The swarm and swarm_hv still run every step, but prediction and speech come from the house, copy and calibration. The Hopfield idea, retrieving from the context itself, is the copy reader. It is the largest single gain measured (+5 points).


Scaling (how she grows)

At a fixed file size, more data helps her slowly: about +0.7 to +0.8 points per doubling of the diet at 1.3 MB. More bytes help steadily: about +1 point per doubling. Her live map% flattening is the house near its ceiling for its size, not a stall.

What moved her, all measured:

Change Effect
Aging removed (it decayed counts every step) +6.3 points
Context fingerprints in every slot (addr 3) 17.4% β†’ 23.1%
In-context copy 23.1% β†’ 28.25%
Calibrated read (SEE + mixed copy) 28.25% β†’ 29.5%, and real probabilities
High-Precision Chooser Head (gated shortlist reranking) +21.4pp on natural shortlists (50.8% β†’ 72.1%), +29.3pp on Cut B shortlists (41.8% β†’ 71.1%), +14.0pp Cut B mixed

Tried and not adopted, because they measured null or negative:

  • a recurrence gate with 4- and 5-token heads
  • count-based context mixing
  • a fold-hashed store
  • a fixed Hopfield store

High-Precision Chooser Upgrade (2026-09-18)

To close the remaining ~20pp choosing gap on candidate shortlists without diluting Zoey's fast, backprop-free core, a calibrated neural reranker was introduced over top-16 candidate shortlists in her native GPT-2 token space (50,257 ids).

Scoring blends causal LM logits with FPHouse vote mass: score(c)=logitLM(c)+Ξ²β‹…ln⁑(1+mc)\text{score}(c) = \text{logit}_{\text{LM}}(c) + \beta \cdot \ln(1 + m_c)

The chooser overrides the house pick if and only if house confidence is contested and the chooser's margin is decisive: (mtop1βˆ’mtop2≀mmax)∧(score(cβˆ—)βˆ’score(chouse)β‰₯Ο„)(m_{\text{top1}} - m_{\text{top2}} \le m_{\text{max}}) \quad \land \quad (\text{score}(c^*) - \text{score}(c_{\text{house}}) \ge \tau) with calibrated parameters $\beta = 1.5$, $m_{\text{max}} = 0.50$, $\tau = 0.80$.

Measured Performance Across Slices

Benchmark Slice N House Pick@1 Chooser Pick@1 Lift Oracle Ceiling Net Flips (W / L) Gated Fire Rate
gold_in_shortlist_no_inject (Cut B) 239 41.84% 71.13% +29.29pp 100.0% +70 (75 wins / 5 losses) 45.6%
natural_gold_in_shortlist (Natural test_pos) 262 50.76% 72.14% +21.37pp 100.0% +56 (68 wins / 12 losses) 40.8%
house_wrong_and_gold_in (Error Recovery) 139 0.00% 53.96% +53.96pp 100.0% +75 (75 wins / 0 losses) 74.8%
mixed_no_inject (Full Cut B, 80% wrong) 500 20.00% 34.00% +14.00pp 47.80% +70 (75 wins / 5 losses) 51.0%
natural_distribution_heldout (Full Natural) 500 26.60% 37.80% +11.20pp 52.40% +56 (68 wins / 12 losses) 47.4%

On Cut B (mixed_no_inject), where 52.2% of rows have no gold token in the top-16 shortlist (oracle ceiling 47.80%), the +14.00pp lift recovers 50.36% of all available oracle headroom.


Live training

  1. Hose. Qwen2.5-14B-Instruct (bf16, 4 prompts per call) generates. Text is re-tokenized to GPT-2 and cut into 2048-token windows, one stream per prompt row. Qwen's raw top-8 distribution is kept at every position where the two tokenizers start on the same byte, about 90% of positions.
  2. Warehouse. Every window is appended to runtime/corpus/zoey_warehouse.jsonl. Qwen's output with its soft targets goes to _stream.jsonl.
  3. Read before write. She predicts each window before learning it; that is the live score. The log reports:
    • pick%: her live reader over the window
    • dense%: the house vote over the window
    • oracle%: how often any head still holds the right token
  4. Write. The window goes into the house (hard counts). Soft-target ingest exists behind a runtime flag and stays off until it is measured.
  5. Checkpoint every 100 steps. The checkpoint records how far into the warehouse the house has eaten, so a restart replays exactly what it had not saved.

Checkpoint

{
  "step": 9300,
  "tokens_seen": 19046400,
  "map_ver": 5,
  "map5": {"ver": 5, "addr": 3, "k1": {}, "k2": {}, "k3": {}, "uni": {}, "writes": 0, "bytes": 1278336},
  "warehouse_consumed": 0,
  "swarm_hv": ["1000 trits"],
  "bf_programs": ["3 programs"],
  "dsl_programs": ["27 programs"]
}

A house from another addr is discarded on load and rebuilt from the warehouse. The programs resume.

Locked encodings

These never change, so old checkpoints and item-memory rows stay bit-identical:

  • token LCG: mul 0x9e3779b97f4a7c15, add 0x6c62272e07bb0142
  • position LCG: mul 0xbf58476d1ce4e5b9, add 0x94d049bb133111eb
  • HV_DIM = 1000

Lineage

  • Schema 4, the ~17 KB soup era: record 21902540, recovery 23354100. ever= reached 14.2989 on the 72B hose.
  • Schema 5, addr 2, on the 14B hose: the 1.31 MB house.
  • Addr 3 (2026-09-11): fingerprinted slots, no aging, the copy reader, the calibrated read, and a 4-wide teacher.

Layout

Path Contents
checkpoints/ Distill + best JSON, LATEST.txt, history
runtime/source/ Trainer, zoey_core (Rust), house, readout, teacher
runtime/logs/ train_online.log, HF upload log
runtime/corpus/ Warehouse jsonl
inference/chat_zoey.py Chat entry
recovery/ Sanitized Linux archive

To run her, load the latest checkpoint JSON, rebuild the swarm, load map5, and start chat_zoey.py. zoey_core must be built with maturin.


Implementation

Piece File Device
House (addr 3) zoey_map5.py CPU
Calibrated read (SEE + copy) zoey_see.py CPU
Readout / generation zoey_readout.py CPU
Trit algebra, fold, VMs zoey_core (Rust / PyO3) CPU
Train loop train_zoey_distill.py CPU + GPU decode
Teacher hose + warehouse zoey_online_teacher.py GPU
High-Precision Chooser cohr_high_precision_chooser.py CPU / GPU
Checkpoint upload zoey_hf_ckpt_loop.py CPU

Created by Dakuwon Moody, Saiyan Corp.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support