YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Qwen3.8-27B DSpark Draft - Stage-1 FineWeb Pretrain (step 104,768, V16)
Boundary checkpoint V16 of the stage-1 token-only (verifier-free) pretraining of a DSpark draft model for Qwen/Qwen3.8-27B, on a multilingual FineWeb mix. This is the final retained checkpoint of the pretrain: the run was deliberately stopped shortly after this boundary (~32.6h of a planned 40h; the few thousand steps trained past it were discarded at the boundary save). It warm-started the FineWeb-stream linear cooldown โ see inference-optimization/qwen38-dspark-cooldown-fineweb, which holds the annealed weights and is the recommended downstream artifact. This repo keeps the pretrain lineage (V14, step 91,672, remains in the file history).
Checkpoint
- global step 104,768 = ~20.6B tokens seen (epoch 0, 10.3% of the 199.06B-token corpus)
- val (V16, 16th consecutive new-best boundary): loss 2.1582, eal 1.6181, position_0_acc 0.3564, accept_rate 0.1416, full_acc 0.2542
- constant LR 1e-3 throughout (no annealing) โ the cooldown repo holds the annealed endpoint
- optimizer / scheduler / training_state intentionally excluded โ downstream phases warm-start with a fresh optimizer (Muon moments are not part of the artifact)
Model
- DSpark draft: 5 transformer layers, hidden 5120, block_size 8, sample_from_anchor, ~0.46B trainable params
- token-only conditioning:
target_layer_ids: [0](fc input [T, 5120] over frozen verifier embeddings; no verifier hidden states) - heads included in weights: markov_head (within-block Markov conditioning) + confidence_head (per-position acceptance prob)
- verifier: Qwen/Qwen3.8-27B (frozen, NOT included - pull from the hub; verifier-owned embed_tokens/lm_head are reconstructed from it)
- loads via speculators
from_pretrained(config.json+model.safetensors;config.pyrecords the model class,val_metrics.jsonthe checkpoint's validation metrics)
Code / environment
- speculators 0.9.0.dev45 @
e7c724808ac8d0fc70740b8bffe29cd7e49a638f(fork of vllm-project/speculators, branchqwen38-27b-dspark-pretraining) - torch 2.13.0, transformers 5.16.1 (vllm 0.30.0 on the training box; not needed for token-only training)
speculators.patch+train_command.txtrecord the exact launch
Training recipe (run.yaml = resolved config, configs/pretrain-fineweb.yaml = source)
- loss:
ce_token(hard-label CE against verifier greedy tokens), EOS<|im_end|> - data mix: 1/2 fineweb-edu (100BT sample, English) + 1/16 x {cmn_Hani, jpn_Jpan, kor_Hang, rus_Cyrl, arb_Arab, spa_Latn, fra_Latn, deu_Latn} from FineWeb-2, token-exact mixing
- Muon, constant lr 1e-3, 51-step warmup, seed 42
- per-rank 24,576-token packed sequence (3,072 anchors x block 8); global 196,608 tokens/step, 8xH200 FSDP
- boundary checkpoint/val every 6,548 steps (~115 min); V16 = the 16th
Metric definitions
eal: greedy expected accepted length per 8-token block incl. the verifier's +1 bonus token (run-based, preserves within-block correlation; ceiling 9.0)accept_rate/accept_len: analytical rejection-sampling quantities (1 - TV overlap)
Lineage / next steps
- cooldown (done): linear LR anneal 1e-3 -> ~0 over 2B tokens, continuing the same FineWeb data stream exactly (
trainer.skip_stepsdata-stream continuation โ no repeated data), warm-started from these weights -> qwen38-dspark-cooldown-fineweb;configs/cooldown-fineweb.yamlin this repo is its source config - fc zero-expansion: fc.weight [5120,5120] -> [5120,46080],
aux_hidden_state_layer_ids[0] -> [0,4,12,20,28,36,44,52,60] - stage-2 hidden-state distillation (fc over verifier hidden states)
- Downloads last month
- 48