How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="livadies/MiniMax-M2.7-Ghetto-MoE-8-Experts", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("livadies/MiniMax-M2.7-Ghetto-MoE-8-Experts", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("livadies/MiniMax-M2.7-Ghetto-MoE-8-Experts", trust_remote_code=True, device_map="auto")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

🩸 MiniMax-M2.7 (229B) -> Ghetto-MoE Edition (8 Experts)

[EN] ⚠️ INFRASTRUCTURE STRESS-TEST KIT

This is a mathematical skeleton (1B parameters) of the giant MiniMax-M2.7 architecture. Created specifically for MLOps engineers to test inference pipelines and routing logic.

Technical Highlights:

  • 8-Expert MoE: Preserved routing logic for Sparse Mixture-of-Experts.
  • RoPE Fixed: Forced linear rotary embeddings to bypass 'default' KeyError.
  • 8:8 GQA: Symmetrical attention heads for consumer-grade GPU support.

[RU] ⚠️ ИНСТРУМЕНТ ДЛЯ ТЕСТИРОВАНИЯ ИНФРАСТРУКТУРЫ

Это математический скелет (1 млрд параметров) гигантской модели MiniMax-M2.7. Создан для отладки инференса и логики роутинга на обычном железе.

Особенности взлома:

  • 8 Экспертов: Сохранена логика переключения экспертов (routing).
  • RoPE Patch: Исправлен баг инициализации через принудительный тип linear.
  • Симметрия GQA: Пропорция голов 8:8 для стабильной работы на старых GPU.

🎧 PROMO DIRECTIVE: RUSSIAN WINTER 26

Слушайте музыку будущего, под которую взламывался этот код:

🔥 Главный трек кванта: «RUSSIAN WINTER 26»

Downloads last month
18
Safetensors
Model size
1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for livadies/MiniMax-M2.7-Ghetto-MoE-8-Experts

Quantized
(112)
this model