YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3.8-27B DSpark Draft - Stage-1 FineWeb Pretrain (step 104,768, V16)

Boundary checkpoint V16 of the stage-1 token-only (verifier-free) pretraining of a DSpark draft model for Qwen/Qwen3.8-27B, on a multilingual FineWeb mix. This is the final retained checkpoint of the pretrain: the run was deliberately stopped shortly after this boundary (~32.6h of a planned 40h; the few thousand steps trained past it were discarded at the boundary save). It warm-started the FineWeb-stream linear cooldown โ€” see inference-optimization/qwen38-dspark-cooldown-fineweb, which holds the annealed weights and is the recommended downstream artifact. This repo keeps the pretrain lineage (V14, step 91,672, remains in the file history).

Checkpoint

  • global step 104,768 = ~20.6B tokens seen (epoch 0, 10.3% of the 199.06B-token corpus)
  • val (V16, 16th consecutive new-best boundary): loss 2.1582, eal 1.6181, position_0_acc 0.3564, accept_rate 0.1416, full_acc 0.2542
  • constant LR 1e-3 throughout (no annealing) โ€” the cooldown repo holds the annealed endpoint
  • optimizer / scheduler / training_state intentionally excluded โ€” downstream phases warm-start with a fresh optimizer (Muon moments are not part of the artifact)

Model

  • DSpark draft: 5 transformer layers, hidden 5120, block_size 8, sample_from_anchor, ~0.46B trainable params
  • token-only conditioning: target_layer_ids: [0] (fc input [T, 5120] over frozen verifier embeddings; no verifier hidden states)
  • heads included in weights: markov_head (within-block Markov conditioning) + confidence_head (per-position acceptance prob)
  • verifier: Qwen/Qwen3.8-27B (frozen, NOT included - pull from the hub; verifier-owned embed_tokens/lm_head are reconstructed from it)
  • loads via speculators from_pretrained (config.json + model.safetensors; config.py records the model class, val_metrics.json the checkpoint's validation metrics)

Code / environment

  • speculators 0.9.0.dev45 @ e7c724808ac8d0fc70740b8bffe29cd7e49a638f (fork of vllm-project/speculators, branch qwen38-27b-dspark-pretraining)
  • torch 2.13.0, transformers 5.16.1 (vllm 0.30.0 on the training box; not needed for token-only training)
  • speculators.patch + train_command.txt record the exact launch

Training recipe (run.yaml = resolved config, configs/pretrain-fineweb.yaml = source)

  • loss: ce_token (hard-label CE against verifier greedy tokens), EOS <|im_end|>
  • data mix: 1/2 fineweb-edu (100BT sample, English) + 1/16 x {cmn_Hani, jpn_Jpan, kor_Hang, rus_Cyrl, arb_Arab, spa_Latn, fra_Latn, deu_Latn} from FineWeb-2, token-exact mixing
  • Muon, constant lr 1e-3, 51-step warmup, seed 42
  • per-rank 24,576-token packed sequence (3,072 anchors x block 8); global 196,608 tokens/step, 8xH200 FSDP
  • boundary checkpoint/val every 6,548 steps (~115 min); V16 = the 16th

Metric definitions

  • eal: greedy expected accepted length per 8-token block incl. the verifier's +1 bonus token (run-based, preserves within-block correlation; ceiling 9.0)
  • accept_rate/accept_len: analytical rejection-sampling quantities (1 - TV overlap)

Lineage / next steps

  1. cooldown (done): linear LR anneal 1e-3 -> ~0 over 2B tokens, continuing the same FineWeb data stream exactly (trainer.skip_steps data-stream continuation โ€” no repeated data), warm-started from these weights -> qwen38-dspark-cooldown-fineweb; configs/cooldown-fineweb.yaml in this repo is its source config
  2. fc zero-expansion: fc.weight [5120,5120] -> [5120,46080], aux_hidden_state_layer_ids [0] -> [0,4,12,20,28,36,44,52,60]
  3. stage-2 hidden-state distillation (fc over verifier hidden states)
Downloads last month
48
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support