VocBulwark speaker encoder

Standalone speaker encoder: turns a raw waveform into a fixed 768-dimensional speaker embedding. It is the conditioning front-end for the companion VocBulwark vocoder โ€” compute an embedding once from a reference clip of a speaker, then pass it to the vocoder to synthesize in that voice. The embedding is also usable on its own for speaker verification / similarity.

Wav2Vec2-based encoder with attentive statistics pooling, trained with a GE2E objective. Self-contained: loads with trust_remote_code=True, no training repo required.

Companion Models

This speaker encoder is part of a set of 6 repositories:

Repo Role
mlr2000/vocoder-large Large vocoder (generates the watermarked audio)
mlr2000/vocoder-large-watermark-detector Watermark detector for the large model
mlr2000/vocoder-large-speaker-encoder Speaker encoder (this repo)
mlr2000/vocoder-small Small vocoder
mlr2000/vocoder-small-watermark-detector Watermark detector for the small model
mlr2000/vocoder-small-speaker-encoder Speaker encoder for the small model

Usage

import torchaudio, torchaudio.functional as AF
from transformers import AutoModel

enc = AutoModel.from_pretrained("mlr2000/vocoder-large-speaker-encoder", trust_remote_code=True).eval()

wav, sr = torchaudio.load("reference.wav")             # [C, T]
wav = wav.mean(0, keepdim=True)                        # mono [1, T]
if sr != enc.config.raw_sample_rate:                   # encoder expects 22.05 kHz
    wav = AF.resample(wav, sr, enc.config.raw_sample_rate)

emb = enc.embed(wav)          # [1, 768]  โ€” feed as speaker_embedding to the vocoder

See example_roundtrip.ipynb in this repo for the full pipeline (reference clip โ†’ embedding โ†’ vocode โ†’ detect watermark).

Notes

  • Input: mono waveform at 22.05 kHz (config.raw_sample_rate). Resample first if your audio differs.
  • Output: [B, 768] L2-comparable speaker embeddings.
  • Pair with the VocBulwark vocoder trained jointly with this encoder โ€” an embedding from a different encoder will not condition the vocoder correctly.
  • This encoder is paired with mlr2000/vocoder-large. For the small vocoder (mlr2000/vocoder-small) use the companion small speaker encoder (mlr2000/vocoder-small-speaker-encoder).
  • Not intended for speaker identification or surveillance of individuals without consent.

Citation

If you use this model, please cite:

@misc{muletta2026,
  title  = {Training a Discriminator-Free Foundation Vocoder 
             with Integrated Audio Watermarking},
  author = {Muletta, Romolo and Deriu, Jan},
  year   = {2026},
  note   = {VT2 Project Report, ZHAW School of Engineering}
}

License

cc-by-4.0. Trained on MLS (CC-BY-4.0) and Common Voice (CC0); builds on BigVGAN (MIT) and wav2vec 2.0 (Apache-2.0). Please retain attribution when redistributing or building on this model.

Downloads last month
11
Safetensors
Model size
18.7M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support