rsi-jev-v5.0-vl-3b

A typed-decision model built by a self-improving loop of AI agents: the first 20 of Qwen/Qwen3.5-4B-Base's 32 layers with the tower fine-tuned and a trained option-scoring head. It answers choice / noul / score questions about a document and up to four images in one forward pass, and speaks the Jev HTTP API.

3.25B parameters: decoder layers 1–20 (2.23B), token embeddings (0.64B), vision tower (0.33B), decision head (0.05B). This repository is self-contained: the tokenizer, the image processor, the embeddings and the vision tower are here, and nothing is read from the base model.

This checkpoint (stages 1–2 seed 17, head stage seed 1):

Decision Index 0.2.1 MMLU-Pro 1k held-out images KoBBQ "unknown" when ambiguous
38.38 0.429 0.829 0.932
pip install "rsi-jev[fast,vision] @ git+https://github.com/Shanghua-Gao/RSI-Jev"
rsi-jev serve v5.0-vl-3b          # or, in Python: from rsijev import Decider; Decider("v5.0-vl-3b")

The loader applies calibration.safetensors, so the probabilities are the calibrated ones this card reports; it builds only the 20 layers the model runs, and reads options with own-token pooling (meta.json: spec.option_pool_own_tokens).

Everything below is the version record from the project repository.


RSI-Jev v5.0-VL 3B

Describes release v5.0-VL · updated 2026-10-02

A System One model doesn't need the whole LLM. v5.0-VL runs the first 20 of Qwen3.5-4B-Base's 32 layers and reads its decision there; the twelve deepest layers, which a language model uses to write the next word, are not in the checkpoint and never run. It takes the same typed questions as before — yes/no, pick-one-of-k, rate-on-a-rubric, about text and up to four images — and returns a calibrated probability for every option in one forward pass.

It is a 3B model: 3.25B parameters — the first 20 decoder layers (2.23B), the token embeddings (0.64B), the vision tower (0.33B) and the decision head (0.05B). The checkpoint is self-contained: everything it runs is in the download, and nothing is fetched from the base model.

Against v4.0-VL (2B) it answers more knowledge questions (MMLU-Pro 0.429 vs 0.385), more image questions (0.829 vs 0.802 on the held-out image set) and far more questions where the right answer is "unknown" (KoBBQ 0.932 vs 0.179), at about 1.45× the latency on a desktop GPU.


1. The task

Unchanged from v4.0-VL: a document (the state), optionally with 1–4 images placed by <image> markers, and one or more typed questions — noul, choice, score — answered with a probability per option in one forward pass. Top-1 is the fraction of decisions whose highest-probability option matches the label; ECE is expected calibration error, lower being better.

2. How the model is built and trained

Base: Qwen3.5-4B-Base, cut at layer 20 of 32. The decision is read from the output of the twentieth decoder layer, after the model's final norm; layers 21–32 are never run. On this base the quality curve flattens by layer 20 (§7): reading at layer 20 scores like reading at layer 32 on the decision suite and the held-out set, and costs 12 fewer layers per token. The option scorer and the confidence head are the same design as v3.0 and v4.0-VL. Images go through the base model's vision tower, frozen, now shipped inside the checkpoint.

Three training stages, each continuing from the one before (specs in data/v5.0-vl_recipe/):

  1. Text, supervised, from the base, 3,000 steps (batch 16, seed 17). A 311,118-question corpus over 68 sources; the lower 8 layers train at a tenth of the learning rate (v2.1's recipe); an LM-retention term is computed at the exit layer, so the cut model keeps what the base knew at layer 20 rather than at layer 32.

  2. Images, supervised, 3,000 steps (half the learning rate). 48,000 questions: 21,754 image questions from the image sources of v4.0-VL and new generators (spatial relations, joint reasoning, recasts, sokoban states), 2,000 new text questions targeting answers of the form "unknown" (KoBBQ and abstention train rows), and 24,246 questions of text replay.

  3. The decision head, under the own-token readout, 600 steps (batch 8, learning rate 1e-5, seed 1). The tower is frozen; only the option scorer trains, on the image stage's own data with its calibration rows held out.

Readout and calibration. Each option's vector is the mean of the tokens of its own block, not counting the separator and key that open it ("\n- option_N:"). Those tokens follow the previous option in a causal model and carry it; pooling them made the 4B line pick the option after the right one on long numbered lists (§4.4). Stages 1 and 2 trained with whole-block pooling; stage 3 retrains the head for the own-token readout, and the confidence head is fitted on top of it (fit seed 0). meta.json records the readout (spec.option_pool_own_tokens: true) and the server reads it.

No RL stage in this release.

3. Checkpoints

Hugging Face shgao/rsi-jev-v5.0-vl-3b
alias v5.0-vl-3b (rsi-jev serve v5.0-vl-3b, Decider("v5.0-vl-3b"))
size 3.25B parameters, 6.2 GB (bf16)
model base 15-benchmark suite held-out (eval_final_v2) suite ECE images download
RSI-Jev-v5.0-VL-3B Qwen3.5-4B-Base, layers 1–20 0.764 0.689 0.050 yes 🤗 download
RSI-Jev-v4.0-VL-2B Qwen3.5-2B-Base 0.756 0.653 0.043 yes 🤗 download
RSI-Jev-v3.0-2B Qwen3.5-2B-Base 0.756 0.649 0.066 no 🤗 download
RSI-Jev-v2.1-2B Qwen3.5-2B-Base 0.736 0.633 0.059 no 🤗 download
RSI-Jev-v1.0-2B Qwen3.5-2B-Base 0.622 0.604 – no 🤗 download

The v4.0-VL to v1.0 rows are from v4.0-VL's record; the v5.0-VL row is its gate run (see the note under §4.1). One seed's checkpoint each.

file holds
tower.safetensors decoder layers 1–20 and the final norm (fine-tuned), token embeddings (the base's)
visual.safetensors the base model's vision tower, unchanged
scorer.safetensors option scorer and confidence head
calibration.json, calibration.safetensors the confidence head's fit (oof_head_scorefloor)
config.json, tokenizer and image-processor files the base's, with num_hidden_layers: 20
meta.json training spec, serving settings, parameter counts

Loading builds only the 20 layers it runs. The package was checked against the full-depth checkpoint it was cut from: bitwise the same probabilities in fp32 (400 questions, 0 changed answers), and the same again when loaded offline with no access to the base model.

The weights are stored in bf16, the precision the server runs them in: served outputs are bitwise the same as from the fp32 training checkpoint (300 questions). Run in fp32, the bf16 weights move probabilities by up to 0.047 and change 4 of 300 answers.

4. Results

All numbers from one scoring pipeline per model; the v4.0-VL column is that release's gate run.

4.1 Against v4.0-VL

v5.0-VL 3B v4.0-VL 2B
MMLU-Pro 1k 0.429 0.385
held-out images (eval_vision_v1) 0.829 0.802
image probes (vision_v3, six) 0.880 0.856
KoBBQ accuracy 0.923 0.538
KoBBQ "unknown" when ambiguous 0.932 0.179
final ECE after calibration 0.042 0.082
suite ECE after calibration 0.050 0.043
15-benchmark suite¹ 0.764 0.756
held-out set (eval_final_v2)¹ 0.689 0.653

¹ Both from the same eval.json aggregation; the gate runs used different scoring drivers, so read these two rows as indicative.

4.2 New image skills

v5.0-VL
spatial relations 0.840
joint reasoning 0.969
sokoban states 0.557

4.3 External: Decision Index 0.2.1

The full suite on this checkpoint, served by rsi-jev serve with nothing cut (all 150,512 requests answered; one RTX PRO 6000): 38.38 (raw 53.33).

area v5.0-VL 3B v4.0-VL 2B
Tools & Automation 52.8 37.7
Language Understanding 44.5 32.6
Retrieval & Classification 43.0 33.3
Knowledge & Reasoning 26.2 19.9
Arts & Human Taste 18.6 11.5
index 38.38 28.31

On the public board (2026-09-28) the best entry at 3.5B served parameters or fewer is Decider 2B at 28.97. Eleven 4B-class entries score below 38.38, among them Intern-Decision-4B (37.81), NeoHorse-Jev-4B (36.75) and Kev 4B (34.64).

4.4 Long numbered option lists

On CLINC150 (151 options keyed option_0…option_150), whole-block pooling made the 4B line pick the option right after the correct one: 39% of the time on 300 items. With own-token pooling and the head retrained for it, the model picks it 2% of the time and accuracy goes from 0.383 to 0.757. On short option lists the two readouts agree within noise (MMLU, HellaSwag, ANLI, WinoGrande, CLadder; 300 items each).

4.5 Calibration

The confidence head is fitted on the retrained head: final-set ECE after calibration 0.042, suite ECE 0.050 (§4.1).

5. Speed and serving

p50, 1 / 8 / 32 questions per request (GB10)
v5.0-VL 3B 43.7 / 124.6 / 338.2 ms

About 1.45× v4.0-VL (2B, all 24 layers) on the GB10, which is compute-bound; on an H100 with CUDA graphs a 4B model cut at 20 runs within 0–13% of the 2B. Serving is unchanged otherwise: inputs up to 32,768 tokens are read whole (a longer question gets a 422), up to 5,120 options per question, structured criteria read as compact JSON.

The API returns a probability and a confidence, and they are different numbers: confidence is the peak statistic Jev documents, (K * p_max - 1) / (K - 1). Route on the probability of the option you intend to act on.

6. Limitations

  • Images are capped at four per request and 1,024 tokens per question. A large photo is scaled down to fit, so small text in a full-page scan can be lost.
  • The vision tower is the base model's, unchanged. Everything this release learned about images is in the text tower and the scorer.
  • Questions longer than 32,768 tokens are refused, not cut (§5).
  • The training data is not in this release. The code and stage specs are here (§7); the data will be published separately.
  • One model, one size, one seed.

7. How it got here

The arms behind this release, in order:

date arm result kept?
2026-09-29 4B base cut at 16, 20, 24, from base, 3k steps, vs a matched 2B suite .744 / .760 / .760 vs .715; the curve flattens by 20 20 taken forward
2026-09-30 images on the layer-20 model, with a text-only control images +.017 to +.34 across image probes, text held; KoBBQ "unknown" fell .828 → .679 not kept
2026-09-30 the same plus 2,000 "unknown"-answer text rows and more abstention images every pre-registered line passed; KoBBQ .899 / .891 kept (this release's weights)
2026-10-01 4B base cut at 12, 13, 28, 32 (depth curve) 12 .681, 13 .712, 28 .758, 32 .761: layer 20 scores like 32 on suite and held-out; MMLU-Pro keeps rising (.422 at 20, .457 at 32) context
2026-10-02 CLINC150 regression traced to option pooling; own-token readout 0.383 → 0.753 on CLINC, short lists unchanged kept (readout)
2026-10-03 the decision head retrained for the own-token readout (600 steps, tower frozen), two seeds final ECE 0.042 (seed 1) and 0.053 (seed 0); Decision Index 38.38 and 38.18 kept (seed 1 = v5.0-VL)

Rerunning it

The training code, the stage specs, the replay source list and the commands are in data/v5.0-vl_recipe/; D is the data directory. The training data itself will be published separately; until then the two stages cannot be rerun from this repository alone.

# stage 1: from the base, read at layer 20, retention at the exit
python scripts/release_train.py --model Qwen/Qwen3.5-4B-Base --seed 17 --root D \
  --spec data/v5.0-vl_recipe/b4-exit20.json --corpus D/rt4-tap16-ret-b --save-dir D/b4-exit20/s17

# stage 2: the image subset, the stage corpus (replay doses protected), then training from stage 1
python scripts/build_vis_subset.py --root D --cap 600 ...      # weights in the recipe README
python scripts/ct_build_corpus.py --parent-corpus D/corpus_hard --steps 3000 --batch 16 \
  --parent-sources "$(cat data/v5.0-vl_recipe/replay_sources.txt)" --protect ... --new D/vis_new/*.jsonl D/text_new/*.jsonl --out D/vis-v4k
python scripts/release_train.py --model Qwen/Qwen3.5-4B-Base --seed 17 --root D \
  --spec data/v5.0-vl_recipe/vis-v4k.json --corpus D/vis-v4k --save-dir D/vis-v4k/s17

# stage 3: the decision head for the own-token readout, tower frozen
RSIJEV_OPTION_POOL_OWN_TOKENS=1 python scripts/fit_head.py --ckpt D/vis-v4k/s17 \
  --corpus D/vis-v4k --dev-corpus D/rt4-tap16-ret-b --out D/headft-B/s17 --seed 1

# the confidence head on it
RSIJEV_OPTION_POOL_OWN_TOKENS=1 python scripts/fit_release_calibration.py --ckpt D/headft-B/s17 --root D \
  --lineage --dev-corpus D/rt4-tap16-ret-b --fit-seed 0 --out D/headft-B/s17

8. Corrections to v4.0-VL's record

v4.0-VL's card is left as published; these are its errata.

  • §1 says three rounds of image training; the shipped lineage has two (vis-ct-1, vis-ct-3).
  • §4.3 "143 pictures" is 143 items on 89 images.
  • Three link texts in it render a literal ../.
  • The matched supervised stage's suite is 0.7579 in the evaluation job (0.7578 in the card).
Downloads last month
103
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for shgao/rsi-jev-v5.0-vl-3b

Finetuned
(207)
this model

Spaces using shgao/rsi-jev-v5.0-vl-3b 2

Collection including shgao/rsi-jev-v5.0-vl-3b