DrumSep (ONNX) β Gridshift mirror
Two models live here: drumsep.onnx (v1, the original weights) and
drumsep_v2.onnx (v2, fine-tuned for electronic drums β what Gridshift loads
from its next release on). See v2 below.
Drum-element separation model (kick / snare / cymbals / toms) in ONNX format,
moved from philippherzig/drumsep-onnx to
Gridshift Studio org for
release-pipeline consistency.
Weights and format are unchanged from the previous repository. The original upstream is inagoy/drumsep, a DemucsHT fine-tune on drum stems.
Usage
Consumed by Gridshift's drum-element separation
feature via the Rust stem-splitter crate. The drumsep_manifest.json in this
repo lists the ONNX model file with SHA-256 for integrity verification.
- Audio format: 44.1 kHz stereo, 1 764 000 samples per inference window
- Stems: kick, snare, cymbals, toms
- Format: ONNX (opset 14)
- File size: ~335 MB
License and attribution
Inherited MIT license from the upstream inagoy/drumsep project. The ONNX
conversion was done from the original HDemucs checkpoint using
torch.onnx.export (see app/rust/drum-separation/convert_drumsep_to_onnx.py
in the Gridshift source tree).
v2 (drumsep_v2.onnx, drumsep_v2_manifest.json)
The v1 weights fine-tuned for electronic drums (808/909-style kicks, claps, processed hats). v1 tends to put kick hits into toms and snare transients into the kick track on electronic material.
- Same shape as v1: four stems, 44.1 kHz stereo, 1 764 000-sample window, opset 14 β a drop-in replacement for the same runtime.
- Training data: synthesized loops β drum one-shots from the Gridshift owner's own library played to MIDI grooves from the Magenta Groove MIDI Dataset (CC BY 4.0) plus generated electronic patterns, with per-stem and bus processing; ~4 500 fine-tuning steps from v1 (L1 plus a transient-weighted and short-window STFT loss).
- Export fix: v1's ONNX had the attention self-reference mask removed by
a manual edit; v2 is exported with the mask intact (
delta == 0instead ofEyeLike) and matches the PyTorch model within 0.4 %. - Held-out synthetic loops, median SDR (v1 β v2): kick 16.5 β 25.1 dB, snare 3.6 β 13.0 dB, cymbals 0.9 β 6.8 dB.
Training and export code: app/Scripts/drumsep/ and
app/rust/drum-separation/convert_drumsep_to_onnx.py in the Gridshift source
tree. License: MIT, inherited from inagoy/drumsep.