rsi-jev-v5.0-vl-3b
A typed-decision model built by a self-improving loop of AI agents: the first 20 of
Qwen/Qwen3.5-4B-Base's 32 layers with the tower fine-tuned and a trained option-scoring head.
It answers choice / noul / score questions about a document and up to four images in one
forward pass, and speaks the
Jev HTTP API.
3.25B parameters: decoder layers 1–20 (2.23B), token embeddings (0.64B), vision tower (0.33B), decision head (0.05B). This repository is self-contained: the tokenizer, the image processor, the embeddings and the vision tower are here, and nothing is read from the base model.
This checkpoint (stages 1–2 seed 17, head stage seed 1):
| Decision Index 0.2.1 | MMLU-Pro 1k | held-out images | KoBBQ "unknown" when ambiguous |
|---|---|---|---|
| 38.38 | 0.429 | 0.829 | 0.932 |
pip install "rsi-jev[fast,vision] @ git+https://github.com/Shanghua-Gao/RSI-Jev"
rsi-jev serve v5.0-vl-3b # or, in Python: from rsijev import Decider; Decider("v5.0-vl-3b")
The loader applies calibration.safetensors, so the probabilities are the calibrated ones this
card reports; it builds only the 20 layers the model runs, and reads options with own-token
pooling (meta.json: spec.option_pool_own_tokens).
Everything below is the version record from the project repository.
RSI-Jev v5.0-VL 3B
Describes release v5.0-VL · updated 2026-10-02
A System One model doesn't need the whole LLM. v5.0-VL runs the first 20 of Qwen3.5-4B-Base's 32 layers and reads its decision there; the twelve deepest layers, which a language model uses to write the next word, are not in the checkpoint and never run. It takes the same typed questions as before — yes/no, pick-one-of-k, rate-on-a-rubric, about text and up to four images — and returns a calibrated probability for every option in one forward pass.
It is a 3B model: 3.25B parameters — the first 20 decoder layers (2.23B), the token embeddings (0.64B), the vision tower (0.33B) and the decision head (0.05B). The checkpoint is self-contained: everything it runs is in the download, and nothing is fetched from the base model.
Against v4.0-VL (2B) it answers more knowledge questions (MMLU-Pro 0.429 vs 0.385), more image questions (0.829 vs 0.802 on the held-out image set) and far more questions where the right answer is "unknown" (KoBBQ 0.932 vs 0.179), at about 1.45× the latency on a desktop GPU.
1. The task
Unchanged from v4.0-VL: a document (the state), optionally with 1–4
images placed by <image> markers, and one or more typed questions — noul, choice,
score — answered with a probability per option in one forward pass. Top-1 is the fraction
of decisions whose highest-probability option matches the label; ECE is expected
calibration error, lower being better.
2. How the model is built and trained
Base: Qwen3.5-4B-Base, cut at layer 20 of 32. The decision is read from the output of the twentieth decoder layer, after the model's final norm; layers 21–32 are never run. On this base the quality curve flattens by layer 20 (§7): reading at layer 20 scores like reading at layer 32 on the decision suite and the held-out set, and costs 12 fewer layers per token. The option scorer and the confidence head are the same design as v3.0 and v4.0-VL. Images go through the base model's vision tower, frozen, now shipped inside the checkpoint.
Three training stages, each continuing from the one before (specs in
data/v5.0-vl_recipe/):
Text, supervised, from the base, 3,000 steps (batch 16, seed 17). A 311,118-question corpus over 68 sources; the lower 8 layers train at a tenth of the learning rate (v2.1's recipe); an LM-retention term is computed at the exit layer, so the cut model keeps what the base knew at layer 20 rather than at layer 32.
Images, supervised, 3,000 steps (half the learning rate). 48,000 questions: 21,754 image questions from the image sources of v4.0-VL and new generators (spatial relations, joint reasoning, recasts, sokoban states), 2,000 new text questions targeting answers of the form "unknown" (KoBBQ and abstention train rows), and 24,246 questions of text replay.
The decision head, under the own-token readout, 600 steps (batch 8, learning rate 1e-5, seed 1). The tower is frozen; only the option scorer trains, on the image stage's own data with its calibration rows held out.
Readout and calibration. Each option's vector is the mean of the tokens of its own block,
not counting the separator and key that open it ("\n- option_N:"). Those tokens follow the
previous option in a causal model and carry it; pooling them made the 4B line pick the option
after the right one on long numbered lists (§4.4). Stages 1 and 2 trained with whole-block
pooling; stage 3 retrains the head for the own-token readout, and the confidence head is fitted
on top of it (fit seed 0). meta.json records the readout (spec.option_pool_own_tokens: true)
and the server reads it.
No RL stage in this release.
3. Checkpoints
| Hugging Face | shgao/rsi-jev-v5.0-vl-3b |
| alias | v5.0-vl-3b (rsi-jev serve v5.0-vl-3b, Decider("v5.0-vl-3b")) |
| size | 3.25B parameters, 6.2 GB (bf16) |
| model | base | 15-benchmark suite | held-out (eval_final_v2) | suite ECE | images | download |
|---|---|---|---|---|---|---|
| RSI-Jev-v5.0-VL-3B | Qwen3.5-4B-Base, layers 1–20 | 0.764 | 0.689 | 0.050 | yes | 🤗 download |
| RSI-Jev-v4.0-VL-2B | Qwen3.5-2B-Base | 0.756 | 0.653 | 0.043 | yes | 🤗 download |
| RSI-Jev-v3.0-2B | Qwen3.5-2B-Base | 0.756 | 0.649 | 0.066 | no | 🤗 download |
| RSI-Jev-v2.1-2B | Qwen3.5-2B-Base | 0.736 | 0.633 | 0.059 | no | 🤗 download |
| RSI-Jev-v1.0-2B | Qwen3.5-2B-Base | 0.622 | 0.604 | – | no | 🤗 download |
The v4.0-VL to v1.0 rows are from v4.0-VL's record; the v5.0-VL row is its gate run (see the note under §4.1). One seed's checkpoint each.
| file | holds |
|---|---|
tower.safetensors |
decoder layers 1–20 and the final norm (fine-tuned), token embeddings (the base's) |
visual.safetensors |
the base model's vision tower, unchanged |
scorer.safetensors |
option scorer and confidence head |
calibration.json, calibration.safetensors |
the confidence head's fit (oof_head_scorefloor) |
config.json, tokenizer and image-processor files |
the base's, with num_hidden_layers: 20 |
meta.json |
training spec, serving settings, parameter counts |
Loading builds only the 20 layers it runs. The package was checked against the full-depth checkpoint it was cut from: bitwise the same probabilities in fp32 (400 questions, 0 changed answers), and the same again when loaded offline with no access to the base model.
The weights are stored in bf16, the precision the server runs them in: served outputs are bitwise the same as from the fp32 training checkpoint (300 questions). Run in fp32, the bf16 weights move probabilities by up to 0.047 and change 4 of 300 answers.
4. Results
All numbers from one scoring pipeline per model; the v4.0-VL column is that release's gate run.
4.1 Against v4.0-VL
| v5.0-VL 3B | v4.0-VL 2B | |
|---|---|---|
| MMLU-Pro 1k | 0.429 | 0.385 |
| held-out images (eval_vision_v1) | 0.829 | 0.802 |
| image probes (vision_v3, six) | 0.880 | 0.856 |
| KoBBQ accuracy | 0.923 | 0.538 |
| KoBBQ "unknown" when ambiguous | 0.932 | 0.179 |
| final ECE after calibration | 0.042 | 0.082 |
| suite ECE after calibration | 0.050 | 0.043 |
| 15-benchmark suite¹ | 0.764 | 0.756 |
| held-out set (eval_final_v2)¹ | 0.689 | 0.653 |
¹ Both from the same eval.json aggregation; the gate runs used different scoring drivers, so
read these two rows as indicative.
4.2 New image skills
| v5.0-VL | |
|---|---|
| spatial relations | 0.840 |
| joint reasoning | 0.969 |
| sokoban states | 0.557 |
4.3 External: Decision Index 0.2.1
The full suite on this checkpoint, served by rsi-jev serve with nothing cut (all 150,512
requests answered; one RTX PRO 6000): 38.38 (raw 53.33).
| area | v5.0-VL 3B | v4.0-VL 2B |
|---|---|---|
| Tools & Automation | 52.8 | 37.7 |
| Language Understanding | 44.5 | 32.6 |
| Retrieval & Classification | 43.0 | 33.3 |
| Knowledge & Reasoning | 26.2 | 19.9 |
| Arts & Human Taste | 18.6 | 11.5 |
| index | 38.38 | 28.31 |
On the public board (2026-09-28) the best entry at 3.5B served parameters or fewer is Decider 2B at 28.97. Eleven 4B-class entries score below 38.38, among them Intern-Decision-4B (37.81), NeoHorse-Jev-4B (36.75) and Kev 4B (34.64).
4.4 Long numbered option lists
On CLINC150 (151 options keyed option_0…option_150), whole-block pooling made the 4B line
pick the option right after the correct one: 39% of the time on 300 items. With own-token pooling
and the head retrained for it, the model picks it 2% of the time and accuracy goes from 0.383 to
0.757. On short option
lists the two readouts agree within noise (MMLU, HellaSwag, ANLI, WinoGrande, CLadder; 300 items
each).
4.5 Calibration
The confidence head is fitted on the retrained head: final-set ECE after calibration 0.042, suite ECE 0.050 (§4.1).
5. Speed and serving
| p50, 1 / 8 / 32 questions per request (GB10) | |
|---|---|
| v5.0-VL 3B | 43.7 / 124.6 / 338.2 ms |
About 1.45× v4.0-VL (2B, all 24 layers) on the GB10, which is compute-bound; on an H100 with CUDA graphs a 4B model cut at 20 runs within 0–13% of the 2B. Serving is unchanged otherwise: inputs up to 32,768 tokens are read whole (a longer question gets a 422), up to 5,120 options per question, structured criteria read as compact JSON.
The API returns a probability and a confidence, and they are different numbers: confidence
is the peak statistic Jev documents, (K * p_max - 1) / (K - 1). Route on the probability of the
option you intend to act on.
6. Limitations
- Images are capped at four per request and 1,024 tokens per question. A large photo is scaled down to fit, so small text in a full-page scan can be lost.
- The vision tower is the base model's, unchanged. Everything this release learned about images is in the text tower and the scorer.
- Questions longer than 32,768 tokens are refused, not cut (§5).
- The training data is not in this release. The code and stage specs are here (§7); the data will be published separately.
- One model, one size, one seed.
7. How it got here
The arms behind this release, in order:
| date | arm | result | kept? |
|---|---|---|---|
| 2026-09-29 | 4B base cut at 16, 20, 24, from base, 3k steps, vs a matched 2B | suite .744 / .760 / .760 vs .715; the curve flattens by 20 | 20 taken forward |
| 2026-09-30 | images on the layer-20 model, with a text-only control | images +.017 to +.34 across image probes, text held; KoBBQ "unknown" fell .828 → .679 | not kept |
| 2026-09-30 | the same plus 2,000 "unknown"-answer text rows and more abstention images | every pre-registered line passed; KoBBQ .899 / .891 | kept (this release's weights) |
| 2026-10-01 | 4B base cut at 12, 13, 28, 32 (depth curve) | 12 .681, 13 .712, 28 .758, 32 .761: layer 20 scores like 32 on suite and held-out; MMLU-Pro keeps rising (.422 at 20, .457 at 32) | context |
| 2026-10-02 | CLINC150 regression traced to option pooling; own-token readout | 0.383 → 0.753 on CLINC, short lists unchanged | kept (readout) |
| 2026-10-03 | the decision head retrained for the own-token readout (600 steps, tower frozen), two seeds | final ECE 0.042 (seed 1) and 0.053 (seed 0); Decision Index 38.38 and 38.18 | kept (seed 1 = v5.0-VL) |
Rerunning it
The training code, the stage specs, the replay source list and the commands are in
data/v5.0-vl_recipe/; D is the data directory. The
training data itself will be published separately; until then the two stages cannot be rerun
from this repository alone.
# stage 1: from the base, read at layer 20, retention at the exit
python scripts/release_train.py --model Qwen/Qwen3.5-4B-Base --seed 17 --root D \
--spec data/v5.0-vl_recipe/b4-exit20.json --corpus D/rt4-tap16-ret-b --save-dir D/b4-exit20/s17
# stage 2: the image subset, the stage corpus (replay doses protected), then training from stage 1
python scripts/build_vis_subset.py --root D --cap 600 ... # weights in the recipe README
python scripts/ct_build_corpus.py --parent-corpus D/corpus_hard --steps 3000 --batch 16 \
--parent-sources "$(cat data/v5.0-vl_recipe/replay_sources.txt)" --protect ... --new D/vis_new/*.jsonl D/text_new/*.jsonl --out D/vis-v4k
python scripts/release_train.py --model Qwen/Qwen3.5-4B-Base --seed 17 --root D \
--spec data/v5.0-vl_recipe/vis-v4k.json --corpus D/vis-v4k --save-dir D/vis-v4k/s17
# stage 3: the decision head for the own-token readout, tower frozen
RSIJEV_OPTION_POOL_OWN_TOKENS=1 python scripts/fit_head.py --ckpt D/vis-v4k/s17 \
--corpus D/vis-v4k --dev-corpus D/rt4-tap16-ret-b --out D/headft-B/s17 --seed 1
# the confidence head on it
RSIJEV_OPTION_POOL_OWN_TOKENS=1 python scripts/fit_release_calibration.py --ckpt D/headft-B/s17 --root D \
--lineage --dev-corpus D/rt4-tap16-ret-b --fit-seed 0 --out D/headft-B/s17
8. Corrections to v4.0-VL's record
v4.0-VL's card is left as published; these are its errata.
- §1 says three rounds of image training; the shipped lineage has two (
vis-ct-1,vis-ct-3). - §4.3 "143 pictures" is 143 items on 89 images.
- Three link texts in it render a literal
../. - The matched supervised stage's suite is 0.7579 in the evaluation job (0.7578 in the card).
- Downloads last month
- 103
Model tree for shgao/rsi-jev-v5.0-vl-3b
Base model
Qwen/Qwen3.5-4B-Base