DrumSep (ONNX) β€” Gridshift mirror

Two models live here: drumsep.onnx (v1, the original weights) and drumsep_v2.onnx (v2, fine-tuned for electronic drums β€” what Gridshift loads from its next release on). See v2 below.

Drum-element separation model (kick / snare / cymbals / toms) in ONNX format, moved from philippherzig/drumsep-onnx to Gridshift Studio org for release-pipeline consistency.

Weights and format are unchanged from the previous repository. The original upstream is inagoy/drumsep, a DemucsHT fine-tune on drum stems.

Usage

Consumed by Gridshift's drum-element separation feature via the Rust stem-splitter crate. The drumsep_manifest.json in this repo lists the ONNX model file with SHA-256 for integrity verification.

  • Audio format: 44.1 kHz stereo, 1 764 000 samples per inference window
  • Stems: kick, snare, cymbals, toms
  • Format: ONNX (opset 14)
  • File size: ~335 MB

License and attribution

Inherited MIT license from the upstream inagoy/drumsep project. The ONNX conversion was done from the original HDemucs checkpoint using torch.onnx.export (see app/rust/drum-separation/convert_drumsep_to_onnx.py in the Gridshift source tree).

v2 (drumsep_v2.onnx, drumsep_v2_manifest.json)

The v1 weights fine-tuned for electronic drums (808/909-style kicks, claps, processed hats). v1 tends to put kick hits into toms and snare transients into the kick track on electronic material.

  • Same shape as v1: four stems, 44.1 kHz stereo, 1 764 000-sample window, opset 14 β€” a drop-in replacement for the same runtime.
  • Training data: synthesized loops β€” drum one-shots from the Gridshift owner's own library played to MIDI grooves from the Magenta Groove MIDI Dataset (CC BY 4.0) plus generated electronic patterns, with per-stem and bus processing; ~4 500 fine-tuning steps from v1 (L1 plus a transient-weighted and short-window STFT loss).
  • Export fix: v1's ONNX had the attention self-reference mask removed by a manual edit; v2 is exported with the mask intact (delta == 0 instead of EyeLike) and matches the PyTorch model within 0.4 %.
  • Held-out synthetic loops, median SDR (v1 β†’ v2): kick 16.5 β†’ 25.1 dB, snare 3.6 β†’ 13.0 dB, cymbals 0.9 β†’ 6.8 dB.

Training and export code: app/Scripts/drumsep/ and app/rust/drum-separation/convert_drumsep_to_onnx.py in the Gridshift source tree. License: MIT, inherited from inagoy/drumsep.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support