Instructions to use mgoeckel/oscar-1-68m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mgoeckel/oscar-1-68m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="mgoeckel/oscar-1-68m")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("mgoeckel/oscar-1-68m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Oscar-1 68M
The 68M rung of the v3 ladder — the recipe that closed the family's transfer gap, at the smallest memory footprint that carries it. The number that matters: 0.7455 typed-decisions test accuracy — above the 144M Julia-1 measured on this host — at ~3.9 ms p50 and 150 MB.
What it is. A decision model: one document (the state) and a set of typed questions in, a
probability distribution per question out, in one forward pass. No text generation — nothing
to parse, nothing to hallucinate. It is a decision head on a JHU CLSP Ettin encoder
(ettin-encoder-68m, fully fine-tuned; ~75M parameters with the head), serving the same
answers contract the hosted Jev API serves, and loads only with the
laya runtime — pip install laya, then laya.load("oscar-1-68m"). The shipped
temperatures (choice 1.050, score 1.061, noul 1.141) were fitted
post-hoc in-distribution on a held-out 400-item split.
Part of the Oscar-1 collection
(all checkpoints + live demo Space).
This version (2026-09-30)
First published release of this member; the recipe lineage runs variants → variants+mix (17m/32m) → chunk_v3 (this ladder).
Read this before relying on these numbers:
- typed-decisions is in-distribution. The training mix contains the same public sources the benchmark's five public suites come from (agnews, mnli, emotion, sst5, banking77) plus the five decision workflows; the test split was held out of training, but the capabilities it measures were trained. Read 0.7455 as what this size learned from the mix, not a general read.
- The sealed 9-suite harness is the transfer check: byte-identical cases at seed 42, labels are the datasets' own (not adjudicated for this harness), 1,240 decisions per participant. With one read and case-clustered resampling, CIs here are wide; treat sub-point gaps as noise.
- As served, confident errors are 25.0% out of domain (decisions at reported confidence ≥ 0.9 that are wrong; the laya anchor: 5.3%). The shipped temperatures were fitted in-distribution; refit on your own decisions before gating anything on confidence.
- No pre-registered release rule for this generation. The family's release rule is registered from the next corpus delta onward (
docs/release-rule.md); this release was decided by reading the results and shipping, which is exactly what the rule is meant to prevent. - mnli 0.508 / sst5 0.358 as served — the v3 recipe's soft-ordinal golds are what keeps these cells alive.
- guardrails 0.896 is one sealed run at 96 decisions — the family's best cell, and the family's widest CI. Do not build policy screens on it yet.
Results (as served: each checkpoint at its own fitted temperature)
| Oscar-1 68M (this checkpoint) | laya anchor | |
|---|---|---|
| typed-decisions test accuracy (400 cases / 2,000 decisions) | 0.7455 | 0.7685 |
| typed-decisions Brier / ECE (as served) | 0.0829 / 0.1171 | 0.0657 / 0.2156 |
| typed-decisions raw accuracy (T = 1.0) | 0.7455 | 0.7685 |
| typed raw Brier / ECE | 0.0851 / 0.1039 | 0.0537 / 0.1301 |
| typed score MAE / within-1 | 0.2776 / 0.965 | 0.2425 / 0.995 |
| typed latency p50 | 3.9 ms | 10.0 ms |
| sysone sealed overall (952 cases / 1,240 decisions) | 0.6629 | 0.7258 |
| sysone agnews | 0.869 | 0.844 |
| sysone banking77 (12) | 0.833 | 0.812 |
| sysone emotion | 0.573 | 0.693 |
| sysone guardrails screens | 0.896 | 0.750 |
| sysone mnli | 0.508 | 0.575 |
| sysone moderation | 0.750 | 0.847 |
| sysone multilingual intent | 0.450 | 0.567 |
| sysone sst5 (5-way sentiment) | 0.358 | 0.375 |
| sysone support triage | 0.734 | 0.927 |
| confident errors out of domain (p ≥ 0.9 and wrong, as served) | 25.0% | 5.3% |
| coverage at ≤ 5% error (share of decisions automatable) | 0.060 | 0.135 |
Latency: typed-decisions readout on an RTX 5060 Ti (bf16); the sealed harness runs the same
checkpoint as served. The laya anchor is our own re-measure on this host, using the same
evaluator that produced the Oscar rows. Its published figures (accuracy 0.766, Tesla T4)
differ from the re-measure within temperature-fitting rounding; its shipped temperatures
include one invalid entry (choice:11+, clamped at load) — accuracy is unaffected.
Paired reads
Paired with oscar-1-150m: sealed overall -1.4 pp [-3.2, +0.5]; 74 decisions right only in oscar-1-150m, 57 only in this checkpoint (exact McNemar p = 0.1619). Level on this harness.
Paired with oscar-1-32m: sealed overall +7.8 pp [+5.4, +10.4]; 77 decisions right only in oscar-1-32m, 174 only in this checkpoint (exact McNemar p ≈ 0). The delta is clear.
Calibration (raw vs served)
Temperature fitting moved typed-decisions accuracy 0.7455 →
0.7455 and Brier 0.0851 → 0.0829 (raw = all three
temperatures at 1.0, set via agent.cfg['temperature'] = [1, 1, 1]). The sealed-harness
runs above are as served only: the harness answers every case through the stock laya
contract with the shipped temperatures.
The family ladder
| member | typed-decisions test acc | sysone sealed overall | confident errors (p ≥ 0.9) | coverage at ≤ 5% error |
|---|---|---|---|---|
oscar-1-17m |
0.6775 | 0.5702 | 19.9% | 0.006 |
oscar-1-32m |
0.7005 | 0.5847 | 20.3% | 0.002 |
| Oscar-1 68M (this checkpoint) | 0.7455 | 0.6629 | 25.0% | 0.060 |
oscar-1-150m |
0.7750 | 0.6766 | 25.4% | 0.019 |
oscar-1-400m |
0.7770 | 0.7000 | 26.6% | 0.002 |
laya-typed-decisions (anchor, measured by us) |
0.7685 | 0.7258 | 5.3% | 0.135 |
All rows measured by us on this host: typed-decisions official test split; sealed 9-suite harness (seed 42). Method and per-decision provenance: see Reproduce.
How it was built
- Backbone: JHU CLSP Ettin
ettin-encoder-68m(bidirectional encoder, fully fine-tuned) + a decision head trained from scratch: 2 transformer layers, an option-marker scorer, an act/escalate head. ~75M parameters together, 149.7 MB shipped on disk. - Recipe (RLCD, the Laya method): the policy reports a full distribution per question; exploration adds zero-mean Gaussian noise to the logits (sigma annealed 0.4 → 0.1); the reward is a strictly proper scoring rule (log + spherical, plus ranked probability score for ordinal
scorequestions), so expected reward is maximised only by honest probabilities. Updates are REINFORCE with a group-mean baseline (GRPO-style), alongside soft cross-entropy against teacher targets (w_ce = 1.0). 4 epochs, 5,300 updates, AdEMAMix, bf16, 2-GPU DDP (0.98 GPU-h per run). - Lineage: trained from the released base encoder (no warm start).
- Corpus: expansion chunk v3 — typed-decisions train mix (customer service, invoice, agent-trace, security, moderation) + public corpora (agnews, mnli, emotion, sst5, banking77) + minted cross-primitive variants + counterfactual repairs + soft-ordinal one-hots + translated multilingual = 84,854 train items (400 held-out calib).
- Calibration: one temperature per primitive fitted post-hoc on a held-out 400-item split (no id overlap with training): choice 1.050, score 1.061, noul 1.141; set
agent.cfg['temperature'] = [1, 1, 1]for the raw readout. - Weights:
model.safetensorssha-256e28aa6e9506a2336…(full digest in results provenance below).
Known limits
- Routing. The memory floor of the v3 ladder (150 MB, 3.9 ms p50): typed 0.7455 clears the 144M Julia-1 measured on this host, and the sealed guardrails cell (0.896) is the family's best. Treat that guardrails number as one wide-CI run until re-measured.
- Out-of-domain calibration is the open problem. As served, decisions at reported confidence ≥ 0.9 are wrong 25.0% of the time on the sealed harness (laya: 5.3%); coverage at a 5% error budget is 0.060 (laya: 0.135). Do not treat confidence as an abstention signal out of distribution; refit or threshold your own data first.
- guardrails measured 0.896 in one sealed run — promising for a built-in prompt-injection screen, one wide-CI read at 96 decisions.
- Ordinal
scoreis the weakest primitive (within-1 0.965, argmax accuracy falls). Keep rubrics to 3-5 levels and read the expected score, not the argmax. - Long states: trained and benchmarked at
max_len = 512(the measured p50 reflects the 512 budget); the ModernBERT encoder itself reads up to ~8,000 tokens — raiseagent.cfg["max_len"]for long states and verify on your own data (the same caveat the laya cards document for their 8k mode). Keep choice sets under ~20 options unlesshead_max_lenis raised too. - English-only. The multilingual intent cell is 0.450 as
served; for non-English states use
laya-multilingual. - A decision model, not an assistant. It cannot answer free-form questions, generate text, or abstain with an "I don't know" option unless you add one to the criteria.
Use
pip install laya
Python 3.10 or newer; CPU inference works, device="cuda" with any modern GPU.
import laya
agent = laya.load("mgoeckel/oscar-1-68m") # downloads and builds
state = {"text": "I was charged twice, please return the money."}
questions = {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": ["refund", "cancel", "information", "other"]},
"urgency": {"type": "score", "instructions": "Rate urgency 0 to 4.",
"criteria": ["routine", "low", "moderate", "high", "critical"]},
"sensitive": {"type": "noul", "instructions": "Is this a fraud/security issue?"},
}
r = agent.predict(state, questions)
print(r["answers"]["intent"]["choice"]) # -> refund
print(r["answers"]["urgency"]["score"]) # -> 0.898 (expected level, 0..1)
print(r["answers"]["sensitive"]["noul"]) # -> probability the answer is true
All questions in a call are answered in one forward pass; latency follows the state's length, not the question count. Answer shape per primitive:
| type | criteria |
answer fields |
|---|---|---|
choice |
2-16 option IDs to descriptions | choice, probabilities, answer_confidence |
score |
ordered rubric levels of 2-16 | score (expected level, 0-1), probabilities, answer_confidence |
noul |
optional true/false descriptions | noul (probability true), confidence = max(p, 1-p) |
The raw JSON is the same answers contract the hosted Jev API and the
laya-demo Space serve — a client
written for either reads Oscar's answers unchanged.
act_probability is present on every answer (act/escalate head, escalate cost 0.5,
wrong-act cost 3.0) but carries no calibrated escalation signal yet — gate decisions on
answer_confidence.
Reproduce
- Sealed harness: sysone-bench v2 orchestrator, dataset version 2.0.0 checksum-verified, seed 42; per-decision predictions, checksums and manifests are archived in
/tmp/rlcd-research/sysone-bench/runs-v2/(regenerable withscripts/run_sysone_gpu.py <name> <checkpoint> <run_id>). - Typed-decisions readout:
encoder_rlcd/eval.pysingle mode over the official test split (400 cases / 2,000 decisions; dataset revision as served through 2026-09). - Result files:
results/oscar-1-68m-sysone.json(sealed),results/ettin-mixed-68m-v3-typeddec.json(typed; raw and as-served blocks),results/oscar-1-paired-sysone.json(paired reads),results/oscar-1-family-sysone.json(as-served confidence metrics),reports/oscar-1-400m-gpu-sweep.mdandreports/oscar-1-68m-150m-gpu-sweep.md(sweep narratives). Pairing method: exact McNemar on discordant decisions + case-clustered bootstrap (10,000 resamples, seed 42). Oscar accuracy is read post-temperature with the per-primitive temperatures inrl_agent_config.json; set all temperatures to 1.0 for the raw readout.
Links
- Collection (all checkpoints + demo): https://huggingface.co/collections/mgoeckel/oscar-1-6abb984de72aa5d5f6978fcc
- Live demo Space: https://huggingface.co/spaces/mgoeckel/oscar-1-demo
- Sibling checkpoints:
oscar-1-17m·oscar-1-32m·oscar-1-150m·oscar-1-400m - Method & runtime: NandhaKishorM/laya ·
pip install laya - Benchmark dataset: https://huggingface.co/datasets/LocalLLaMA/typed-decisions
- Companion models:
convaiinnovations/laya-typed-decisions·SupersonicLabs/Julia-1·openjev/openjev
Apache 2.0 · Oscar-1 (mgoeckel) · Ettin encoders MIT (JHU CLSP) · RLCD method & Laya runtime Apache-2.0 (Convai Innovations)
- Downloads last month
- 29
Model tree for mgoeckel/oscar-1-68m
Base model
jhu-clsp/ettin-encoder-68mDataset used to train mgoeckel/oscar-1-68m
Collection including mgoeckel/oscar-1-68m
Evaluation results
- accuracy on typed-decisions (official test split, measured by us)self-reported0.746
- brier_score on typed-decisions (official test split, measured by us)self-reported0.083
- expected_calibration_error on typed-decisions (official test split, measured by us)self-reported0.117
- mean_absolute_error on typed-decisions (official test split, measured by us)self-reported0.278
- sysone sealed overall (9-suite harness, 1,240 decisions, seed 42) on typed-decisions (official test split, measured by us)self-reported0.663