Reference wav not supported?

#1
by racaelin - opened

Having trouble getting this to work with a reference wav.

Running the example command:

VIBEVOICE_VOICE_AUDIO=input.wav
./build/bin/crispasr --tts "Hello, how are you today?"
-m vibevoice-1.5b-tts-q4_k.gguf
--tts-output output.wav
crispasr 0.6.0 (git a471a3c6, Release) [backends: cpu]
crispasr: auto-detected backend 'vibevoice' from '/cachecrispasr/vibevoice-1.5b-tts-q4_k.gguf'
vibevoice: d_lm=1536, layers=28, heads=12/2, ffn=8960, vocab=151936
vibevoice: vae_acoustic=64, vae_semantic=128, downsample=3200x
vibevoice: loaded 1204 tensors (backend: CPU)
crispasr[vibevoice-tts]: no voice prompt resolved (pass --voice

When adding the --voice path argument suggested from error output, it still errors out:

gguf_init_from_file_ptr: invalid magic characters: 'RIFF', expected 'GGUF'
vibevoice-voice: failed to load tensor metadata from 'input.wav'
vibevoice: failed to load voice prompt 'input.wav'
crispasr[vibevoice-tts]: voice 'input.wav' could not be loaded; refusing to synthesise without a voice prompt.
crispasr: error: TTS synthesis failed

Seems like the current git version expects precomputed embeddings instead of .wav?

Owner

Fixed in commit fbda63b6 (2026-05-01) 'fix(vibevoice-tts): fix 1.5B WAV-clone prompt + normalization + silence trim', shipped in v0.6.2 and later. You're on v0.6.0 (a471a3c6). --voice <path>.wav works on v0.6.3+.

cstr changed discussion status to closed

Sign up or log in to comment