Instructions to use maya-research/maya1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use maya-research/maya1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-to-speech", model="maya-research/maya1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("maya-research/maya1") model = AutoModelForCausalLM.from_pretrained("maya-research/maya1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Maya-1
Maya-1 is an open-weight text-to-speech model from Maya Research that combines natural-language voice descriptions with inline expressive controls. It generates speech as discrete audio tokens, which are decoded into a 24 kHz waveform. Developers can describe a voice and its delivery in text instead of choosing only from a fixed speaker menu.
Demos
Energetic Female Event HostVoice descriptionFemale, in her 30s with an American accent and is an event host, energetic, clear diction |
Calm Male NarratorVoice descriptionMale, late 20s, neutral American, warm baritone, calm pacing |
Model at a glance
| Property | Released model |
|---|---|
| Repository | maya-research/maya1 |
| Task | Text-to-speech with description-based voice design and expressive controls |
| Parameters | Approximately 3.30 billion; Hub tensor metadata reports 3,300,928,512 parameters |
| Model implementation | Decoder-only LlamaForCausalLM |
| Audio representation | SNAC discrete audio codes |
| Decoder used in the reference implementation | hubertsiuzdak/snac_24khz |
| Native output | 24 kHz mono audio |
| Released weight precision | BF16 |
| Published expressive range | More than 20 emotions and vocal styles, as described in the release |
| Distribution | Hugging Face checkpoint, tokenizer and reference inference code |
| Model license declaration | Apache 2.0 in repository metadata; see component-license notes below |
This card documents the Maya-1 checkpoint. Other Maya Research models have separate documentation.
1. Purpose and design
Design objective
Maya-1 separates what should be spoken from how it should sound, while bringing both into the model's generation process. An utterance provides the words, a natural-language description guides the overall voice, and inline markers request local expression. Together, these inputs condition the acoustic-token sequence that becomes the output waveform.
The design connects three parts: a controllable text interface, training examples that pair written direction with audio, and an inference path that turns generated acoustic tokens into playable speech. The sections below explain these connections and the decisions behind them.
Technical contributions and novelty
Maya-1 brings descriptive voice direction, local expressive control and acoustic-token generation together in one trained, open-weight model. Its model-level contribution is the way these capabilities are connected in the conditioning interface, training workflow and released inference path.
- A voice specified through language: the model receives a description of the desired voice alongside the utterance. This makes age, timbre, accent, pacing and delivery part of the generative input rather than restricting the interface to selection from a predefined speaker list.
- Voice identity and expression controlled together: description-level guidance establishes the overall voice, while inline markers request local delivery changes or vocalizations. A passage can retain its voice direction while requesting laughter, whispering or other expressive behavior at relevant points.
- Supervision tied to the conditioning interface: curated recordings, human-reviewed voice descriptions and expressive annotations provide the training material for connecting written direction with acoustic output. The data workflow and conditioning-format experiments are described below.
- Conditioning format development: exploration of colon-delimited descriptions, attribute lists and key-value tags led to the selected description wrapper. The reported qualitative observations explain the choice: preserve flexible natural-language direction while reducing format drift, spoken descriptions and sensitivity to rigid attribute syntax.
- A complete path from conditioned text to playable audio: the causal transformer predicts SNAC acoustic codes, the decoding path reconstructs the hierarchical codec streams, and the streaming path emits audio incrementally. The model exposes the resulting capability through downloadable weights and inference resources.
Maya-1 builds on a Llama-style transformer and the SNAC codec, connecting these components through its conditioning interface, training workflow and inference implementation. The contribution described here is this model-specific design and integration.
Key design choices
The following choices explain the technical purpose of this design and the practical considerations associated with it.
1. Use a natural-language description to condition the voice.
- Design: provide a voice description alongside the utterance. The description specifies the desired characteristics; the utterance supplies the spoken content. Both enter the model's input sequence.
- Rationale: developers can request combinations of accent, pitch, timbre, pace and delivery in ordinary language, without selecting only from a fixed speaker menu or managing a separate numeric control for every attribute.
- Usage consideration: descriptions work best as clear, compatible requests. Ambiguous or conflicting attributes can reduce consistency.
2. Combine overall voice guidance with local expressive markers.
- Design: combine description-level voice guidance with inline tokens such as
<whisper>,<laugh>and<sigh>. - Rationale: overall voice characteristics and moment-to-moment expression are different control needs. Inline markers let a passage request changes or nonverbal sounds near the relevant spoken content, complementing the global description.
- Usage consideration: inline tokens request delivery or vocal events; their placement does not specify an exact timestamp, duration or intensity.
3. Model speech as a sequence of discrete audio tokens.
- Design: predict SNAC acoustic codes autoregressively, then reconstruct the waveform with the SNAC decoder. The reference implementation groups seven generated audio tokens into each frame and reconstructs three hierarchical streams.
- Rationale: a compact acoustic-token representation makes audio generation a sequence-prediction task while delegating waveform reconstruction to the codec decoder. The transformer does not have to predict individual waveform samples directly.
- Engineering consideration: token ordering, codec compatibility and decoding boundaries are important to output quality.
4. Use a Llama-style decoder-only transformer as the sequence model.
- Design: use an approximately 3.30B-parameter decoder-only transformer with 28 layers, 24 attention heads and eight key-value heads.
- Rationale: causal attention lets each next audio-token prediction use the preceding voice description, utterance and generated sequence. The 24-query-head/eight-key-value-head configuration shares key-value heads across groups of queries, reducing cached key-value heads relative to an otherwise equivalent configuration with 24 key-value heads.
- Engineering consideration: deployment memory includes the transformer, key-value cache, audio decoder and framework overhead. Hardware suitability should be established for the intended input lengths and concurrency.
5. Decode audio incrementally rather than wait for the whole utterance.
- Design: accumulate generated audio tokens, decode overlapping four-frame windows and emit waveform chunks.
- Rationale: downstream playback can begin while the rest of the utterance is still being generated. This serves interfaces where the delay before the first audible response matters, rather than optimizing only for when a complete audio file becomes available.
- Engineering consideration: incremental decoding trades an initial buffering requirement and repeated decoding work for earlier output. End-to-end latency also includes transport and playback.
6. Distribute the weights and inference interface for external use.
- Release choice: distribute the weights, tokenizer, configuration and inference resources through Hugging Face, with an Apache 2.0 model-license declaration.
- Rationale: developers can inspect, evaluate and integrate the checkpoint on their own infrastructure rather than depend exclusively on a single hosted API, giving them control over the serving environment and application integration.
- Deployment consideration: self-hosting includes responsibility for compute, maintenance, safeguards and application integration.
2. Inputs and expressive controls
Maya-1 takes two conceptual inputs:
- Voice description: the voice and delivery characteristics to request.
- Utterance text: the words to synthesize, optionally containing expressive control tokens.
For example:
Voice description:
A warm, low-pitched narrator with clear diction and an unhurried
conversational pace.
Utterance:
The experiment worked. <sigh> We can finally take a breath.
The description wrapper documented for the model is:
<description="VOICE DESCRIPTION"> UTTERANCE TEXT
This illustrates the input content, not the entire token sequence required by an inference backend. Model-specific control tokens and generation boundaries must also be applied. The Transformers README example and the separate vLLM example currently use different prompt-construction paths; they should not be assumed interchangeable without testing.
Expressive range and tokenizer inventory
The published release describes support for more than 20 emotions and vocal styles through voice-description conditioning and inline controls. This expressive-range description is separate from the inventory of registered tokens below.
The released tokenizer declares the following emotion and vocal-style tokens:
<angry> <appalled> <chuckle> <cry>
<curious> <disappointed> <excited> <exhale>
<gasp> <giggle> <gulp> <laugh>
<laugh_harder> <mischievous> <sarcastic> <scream>
<sigh> <sing> <snort> <whisper>
The inventory includes emotional delivery, nonverbal vocalizations and other vocal styles. Results depend on the text, voice description, token placement and sampling settings; examples should be auditioned for the intended use.
The released tokenizer is the source for this inventory.
3. Architecture and audio representation
Transformer
The published configuration specifies:
| Configuration field | Value |
|---|---|
| Architecture | LlamaForCausalLM |
| Transformer layers | 28 |
| Hidden dimension | 3,072 |
| Feed-forward intermediate dimension | 8,192 |
| Attention heads | 24 |
| Key-value heads | 8 |
| Head dimension | 128 |
| Vocabulary size | 156,960 |
| Activation | SiLU |
| Weight tying | Enabled |
The implementation uses a Llama-style decoder-only transformer. With 24 attention heads and eight key-value heads, groups of three query heads share a key-value head. This reduces the number of key-value heads cached during generation relative to an otherwise equivalent 24-key-value-head configuration.
The configuration contains max_position_embeddings=131072, while the published streaming example configures an 8,192-token model length. Usable generation length and long-form voice consistency should be evaluated separately from the configured position limit.
Audio-token generation
The reference inference path can be summarized as:
Voice description + utterance + control framing
β
Transformer token generation
β
Seven-token SNAC frames
β
Three hierarchical codec streams
β
SNAC decoder
β
24 kHz mono waveform
The released decoding implementation groups seven generated codec tokens into a frame. Those tokens are unpacked into three hierarchical streams with one, two and four codes per frame respectively. The reference code uses audio-token IDs 128266 through 156937, with 128257 and 128258 as speech-start and speech-end control IDs.
The transformer produces the speech-token sequence, and the upstream SNAC decoder reconstructs the waveform.
Streaming
The vLLM reference implementation buffers generated audio tokens, decodes overlapping four-frame windows and emits waveform chunks. It demonstrates incremental decoding rather than requiring the complete utterance to finish first.
Time to first audio, sustained throughput and end-to-end playback latency depend on hardware, precision, batching, prompt length, decoding and transport. A useful deployment benchmark reports these separately from the duration of each emitted audio chunk.
The example emits 24 kHz mono, 16-bit PCM, equivalent to 384 kbit/s before transport overhead. This waveform output format is separate from the compressed acoustic-token representation inside the model.
Training and data preparation
Maya-1's data-curation priorities included an India-first focus.
Maya-1's published development account describes speech pretraining for broad acoustic coverage and natural coarticulation, followed by supervised fine-tuning on curated studio recordings. The fine-tuning examples paired audio with human-reviewed voice descriptions and expressive annotations, with variations in accent, character and delivery.
This supervision connects the conditioning interface to the training task: the voice description supplies guidance on how the utterance should sound, while the target audio provides the acoustic realization. It complements the architecture and expressive controls described above.
The reported data-processing workflow included:
- Audio standardization: resampling recordings to 24 kHz mono and normalizing loudness to -23 LUFS.
- Silence and duration processing: trimming silence with voice-activity detection and applying duration bounds of 1 to 14 seconds.
- Phrase alignment: using Montreal Forced Aligner to establish clean phrase boundaries.
- Text deduplication: using MinHash-LSH to identify repeated or closely similar text.
- Audio deduplication: using Chromaprint to identify duplicate audio.
- Acoustic-token preparation: encoding the audio with SNAC and packing the representation into seven-token frames for model development.
Conditioning format experiments
The original development account describes exploration of four formats for combining voice direction with utterance text. The qualitative observations concerned formatting stability, flexibility and whether the description remained a control instruction rather than being spoken.
| Conditioning format | Reported observation |
|---|---|
| Description followed by a colon and utterance | Format drift; the model sometimes spoke the description |
| Angle-bracket attribute list | Rigid representation and weaker generalization |
| Separate key-value tags | Longer markup and sensitivity to input errors |
| Natural-language description wrapper | Selected for flexible voice descriptions and more robust formatting |
The selected content format was:
<description="VOICE DESCRIPTION"> UTTERANCE TEXT
The decision retained free-form voice direction within an explicit wrapper while keeping the utterance outside that wrapper. The format comparison reports qualitative development observations rather than numerical ablation scores.
Deployment and performance
The published release describes single-GPU deployment, vLLM integration and incremental audio output, with the following infrastructure and streaming features.
Published deployment features
- Runs on a single GPU.
- vLLM integration for scale.
- Automatic prefix caching for efficiency.
- 24 kHz audio output.
- WebAudio-compatible output for browser playback.
Streaming deployment features described in the release
- Automatic Prefix Caching (APC) for repeated voice descriptions.
- WebAudio ring-buffer integration.
- Multi-GPU scaling support.
- Sub-100ms latency for real-time applications.
Sub-100ms latency is reported in the release, not established by a new benchmark in this card. Achieved latency depends on hardware, serving configuration, prompt length, decoding and transport. Measure time to first audio separately from the duration of each emitted audio chunk; the two are not interchangeable.
4. Intended use and limitations
Potential applications include expressive narration, fictional character dialogue, assistive reading and conversational interfaces. These are intended uses, not validation that the model satisfies every application's requirements.
Important limitations:
- Language and accent variation: pronunciation, intelligibility and expressive control can vary by language and accent. Evaluate representative inputs from the intended language, including code-switching where relevant, before deployment.
- Description following: requested accents, timbres and styles are approximate controls. Conflicting or unusually detailed descriptions may not be followed consistently.
- Pronunciation and content fidelity: names, numbers, abbreviations and unusual text should be checked in generated audio. Generative output can omit, repeat or mispronounce content.
- Long-form consistency: the configured token capacity does not establish consistent speaker identity, pacing or intelligibility over long recordings.
- Control reliability: expressive-token behavior varies with the text, voice description, placement and sampling settings.
- Deployment: checkpoint availability does not supply authentication, request limits, abuse monitoring, consent verification or a production service-level commitment.
- Evaluation coverage: applications should assess performance across their relevant accents, speaking styles and user groups.
Responsible deployment
Clearly disclose synthetic voice where appropriate. Do not use generated audio for deceptive impersonation, fraud, fabricated evidence or unauthorized representations of a real person's statements. Applications should address consent, access controls, abuse handling and retention of submitted text and audio.
Accessibility and other consequential applications should include human review and an appropriate fallback. Application-level safeguards must be implemented by the deployer.
5. Inference resources and reproducibility
Original Transformers inference example
This is the full inference program from the published Maya-1 card, retaining its prompt construction, generation, SNAC reconstruction and WAV output. One tokenization correction is included: add_special_tokens=False avoids prepending an additional beginning-of-sequence token to a prompt that already contains one.
Use a CUDA environment with BF16 support and sufficient memory for the model and audio decoder. Install the example's dependencies:
pip install torch transformers accelerate snac soundfile numpy
Save the following program as maya1_example.py, then run python maya1_example.py. The original example downloads the model and decoder if they are not already cached and writes output.wav.
#!/usr/bin/env python3
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from snac import SNAC
import soundfile as sf
import numpy as np
CODE_START_TOKEN_ID = 128257
CODE_END_TOKEN_ID = 128258
CODE_TOKEN_OFFSET = 128266
SNAC_MIN_ID = 128266
SNAC_MAX_ID = 156937
SNAC_TOKENS_PER_FRAME = 7
SOH_ID = 128259
EOH_ID = 128260
SOA_ID = 128261
BOS_ID = 128000
TEXT_EOT_ID = 128009
def build_prompt(tokenizer, description: str, text: str) -> str:
"""Build formatted prompt for Maya1."""
soh_token = tokenizer.decode([SOH_ID])
eoh_token = tokenizer.decode([EOH_ID])
soa_token = tokenizer.decode([SOA_ID])
sos_token = tokenizer.decode([CODE_START_TOKEN_ID])
eot_token = tokenizer.decode([TEXT_EOT_ID])
bos_token = tokenizer.bos_token
formatted_text = f'<description="{description}"> {text}'
prompt = (
soh_token + bos_token + formatted_text + eot_token +
eoh_token + soa_token + sos_token
)
return prompt
def extract_snac_codes(token_ids: list) -> list:
"""Extract SNAC codes from generated tokens."""
try:
eos_idx = token_ids.index(CODE_END_TOKEN_ID)
except ValueError:
eos_idx = len(token_ids)
snac_codes = [
token_id for token_id in token_ids[:eos_idx]
if SNAC_MIN_ID <= token_id <= SNAC_MAX_ID
]
return snac_codes
def unpack_snac_from_7(snac_tokens: list) -> list:
"""Unpack 7-token SNAC frames to 3 hierarchical levels."""
if snac_tokens and snac_tokens[-1] == CODE_END_TOKEN_ID:
snac_tokens = snac_tokens[:-1]
frames = len(snac_tokens) // SNAC_TOKENS_PER_FRAME
snac_tokens = snac_tokens[:frames * SNAC_TOKENS_PER_FRAME]
if frames == 0:
return [[], [], []]
l1, l2, l3 = [], [], []
for i in range(frames):
slots = snac_tokens[i*7:(i+1)*7]
l1.append((slots[0] - CODE_TOKEN_OFFSET) % 4096)
l2.extend([
(slots[1] - CODE_TOKEN_OFFSET) % 4096,
(slots[4] - CODE_TOKEN_OFFSET) % 4096,
])
l3.extend([
(slots[2] - CODE_TOKEN_OFFSET) % 4096,
(slots[3] - CODE_TOKEN_OFFSET) % 4096,
(slots[5] - CODE_TOKEN_OFFSET) % 4096,
(slots[6] - CODE_TOKEN_OFFSET) % 4096,
])
return [l1, l2, l3]
def main():
# Load the best open source voice AI model
print("\n[1/3] Loading Maya1 model...")
model = AutoModelForCausalLM.from_pretrained(
"maya-research/maya1",
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
"maya-research/maya1",
trust_remote_code=True
)
print(f"Model loaded: {len(tokenizer)} tokens in vocabulary")
# Load SNAC audio decoder (24kHz)
print("\n[2/3] Loading SNAC audio decoder...")
snac_model = SNAC.from_pretrained("hubertsiuzdak/snac_24khz").eval()
if torch.cuda.is_available():
snac_model = snac_model.to("cuda")
print("SNAC decoder loaded")
# Design your voice with natural language
description = "Realistic male voice in the 30s age with american accent. Normal pitch, warm timbre, conversational pacing."
text = "Hello! This is Maya1 <laugh_harder> the best open source voice AI model with emotions."
print("\n[3/3] Generating speech...")
print(f"Description: {description}")
print(f"Text: {text}")
# Create prompt with proper formatting
prompt = build_prompt(tokenizer, description, text)
# Debug: Show prompt details
print(f"\nPrompt preview (first 200 chars):")
print(f" {repr(prompt[:200])}")
print(f" Prompt length: {len(prompt)} chars")
# Generate emotional speech
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
print(f" Input token count: {inputs['input_ids'].shape[1]} tokens")
if torch.cuda.is_available():
inputs = {k: v.to("cuda") for k, v in inputs.items()}
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=2048, # Increase to let model finish naturally
min_new_tokens=28, # At least 4 SNAC frames
temperature=0.4,
top_p=0.9,
repetition_penalty=1.1, # Prevent loops
do_sample=True,
eos_token_id=CODE_END_TOKEN_ID, # Stop at end of speech token
pad_token_id=tokenizer.pad_token_id,
)
# Extract generated tokens (everything after the input prompt)
generated_ids = outputs[0, inputs['input_ids'].shape[1]:].tolist()
print(f"Generated {len(generated_ids)} tokens")
# Debug: Check what tokens we got
print(f" First 20 tokens: {generated_ids[:20]}")
print(f" Last 20 tokens: {generated_ids[-20:]}")
# Check if EOS was generated
if CODE_END_TOKEN_ID in generated_ids:
eos_position = generated_ids.index(CODE_END_TOKEN_ID)
print(f" EOS token found at position {eos_position}/{len(generated_ids)}")
# Extract SNAC audio tokens
snac_tokens = extract_snac_codes(generated_ids)
print(f"Extracted {len(snac_tokens)} SNAC tokens")
# Debug: Analyze token types
snac_count = sum(1 for t in generated_ids if SNAC_MIN_ID <= t <= SNAC_MAX_ID)
other_count = sum(1 for t in generated_ids if t < SNAC_MIN_ID or t > SNAC_MAX_ID)
print(f" SNAC tokens in output: {snac_count}")
print(f" Other tokens in output: {other_count}")
# Check for SOS token
if CODE_START_TOKEN_ID in generated_ids:
sos_pos = generated_ids.index(CODE_START_TOKEN_ID)
print(f" SOS token at position: {sos_pos}")
else:
print(f" No SOS token found in generated output!")
if len(snac_tokens) < 7:
print("Error: Not enough SNAC tokens generated")
return
# Unpack SNAC tokens to 3 hierarchical levels
levels = unpack_snac_from_7(snac_tokens)
frames = len(levels[0])
print(f"Unpacked to {frames} frames")
print(f" L1: {len(levels[0])} codes")
print(f" L2: {len(levels[1])} codes")
print(f" L3: {len(levels[2])} codes")
# Convert to tensors
device = "cuda" if torch.cuda.is_available() else "cpu"
codes_tensor = [
torch.tensor(level, dtype=torch.long, device=device).unsqueeze(0)
for level in levels
]
# Generate final audio with SNAC decoder
print("\n[4/4] Decoding to audio...")
with torch.inference_mode():
z_q = snac_model.quantizer.from_codes(codes_tensor)
audio = snac_model.decoder(z_q)[0, 0].cpu().numpy()
# Trim warmup samples (first 2048 samples)
if len(audio) > 2048:
audio = audio[2048:]
duration_sec = len(audio) / 24000
print(f"Audio generated: {len(audio)} samples ({duration_sec:.2f}s)")
# Save your emotional voice output
output_file = "output.wav"
sf.write(output_file, audio, 24000)
print(f"\nVoice generated successfully!")
if __name__ == "__main__":
main()
The restored program has passed a Python syntax check. No new GPU inference or generated-audio check was performed for this documentation revision.
Optional companion example
The optional companion example generates a 24 kHz mono WAV, pins model and decoder revisions, checks prompt boundaries and reconstructs the three codec streams. It uses cached weights unless downloading is enabled.
After preparing the environment described in the companion instructions, run:
python maya1_inference.py \
--description 'A warm narrator with clear diction.' \
--text 'Hello. This is a short voice generation example.' \
--output maya1-example.wav \
--allow-download
This adaptation has passed eight protocol unit tests, Python compilation and command-line checks. Model loading and generated audio have not yet been tested end to end in the documented environment. The verification record separates completed checks from the remaining inference and listening checks. It is a batch example, not a streaming server or latency benchmark.
Published reference paths
The repository contains a Transformers generation example in the existing README at the inspected revision and a vLLM streaming reference.
The streaming script is a reference implementation, not an unchanged copy-and-run production service. It contains a developer-local checkpoint path that must be replaced with the repository ID or a valid local checkpoint. Its generic chat-template prompt path differs from the explicit Transformers framing used in the companion example. Validate the chosen path before deployment; the two are not asserted to be equivalent.
The released generation_config.json specifies temperature 0.6, while the examples use 0.4. These are different settings, not one canonical default. The appropriate values depend on the application and should be reported with evaluations.
For reproducible use, record the checkpoint revision, tokenizer revision, decoder version, serving-library versions, hardware, precision, generation settings, input length and relevant latency boundary. This card describes revision:
21c682a0afef8c13a89b2512733c8bf5f0c52eb7
6. Contributors
Maya-1 was developed at Maya Research. Dheemanth Reddy led Maya-1's model design and data curation and contributed hands-on to its development, while Bharath Kumar Kakumani contributed to model development, data preparation, compute coordination, and checkpoint release.
Dheemanth Reddy
Model-design and data-curation lead; model co-developer
Dheemanth's model-design work connected the goal of controllable voice generation with the conditioning interface and text-to-audio architecture. The following areas explain the scope of that design leadership and the purpose of the resulting design choices:
- Model development: contributing directly to building Maya-1, alongside his model-design and data-curation responsibilities.
- Voice-description conditioning: bringing natural-language voice direction into the model's input alongside the words to be spoken. In Maya-1, the description guides voice characteristics and delivery, while the utterance supplies the spoken content. This gives developers a way to request combinations of vocal qualities through text, rather than relying only on a fixed menu of speakers.
- Overall voice guidance and local expression: combining description-level guidance with inline expressive controls. The description supplies the overall voice direction; markers such as
<whisper>,<laugh>and<sigh>request delivery changes or vocalizations near particular passages. These serve different control needs within the same utterance, without treating every expressive change as a new speaker selection. - Text-to-audio architecture: bringing the conditioning interface together with a causal transformer that predicts SNAC acoustic tokens. This connects voice instructions and spoken content to a sequence of audio codes; the SNAC decoder then reconstructs the waveform. The design separates acoustic-token prediction from waveform reconstruction while conditioning acoustic-token generation on the description, utterance and expressive cues within one input sequence.
- Data curation: leading Maya-1's data-curation work with an India-first focus, alongside the model-design work, with Bharath supporting data preparation. The training section describes the team's use of voice descriptions, expressive annotations and audio preparation to connect conditioning inputs with the target voice output.
- Open release: contributing to making Maya-1 available under Apache 2.0. Releasing the weights and accompanying model resources allows developers to inspect, adapt and integrate the model in their own environments.
Bharath Kumar Kakumani
Model development, data preparation, compute and checkpoint release
Bharath co-developed Maya-1 with Dheemanth. His contributions spanned the model's technical development, preparation of development data, coordination of compute resources and publication of the checkpoint:
- Model development: contributing to Maya-1's technical implementation alongside Dheemanth's model-design and data-curation leadership. This was part of the shared development effort that brought voice-description conditioning, expressive controls and acoustic-token generation together in the released model.
- Data preparation: supporting preparation of the data used for model development, working alongside Dheemanth's curation work. This contribution helped make curated material usable within the team's model-development workflow.
- GPU-resource coordination: helping coordinate access to the GPU resources used for model development. His contribution supported the team's access to the compute needed to develop Maya-1.
- Checkpoint publication: handling the upload of Maya-1's model checkpoint to Hugging Face as part of the public release, making the released weights available for developers to download, evaluate and integrate into their own applications.
7. License and acknowledgments
The model repository declares Apache 2.0 in its metadata. The streaming example separately carries an MIT notice. Preserve the applicable notices for the model, example code and dependencies when redistributing or adapting them.
Maya-1's released implementation uses or builds on:
- SNAC for discrete audio representation and decoding.
- The Llama-style architecture implemented through Transformers.
- vLLM in the published streaming reference.
8. Citation
Cite the model repository and the checkpoint revision used:
@misc{maya1_model,
title = {Maya-1},
author = {{Maya Research}},
year = {2025},
howpublished = {Hugging Face model repository},
url = {https://huggingface.co/maya-research/maya1}
}
- Downloads last month
- 6,492