A newer version of this model is available: IvmeLabs/Ivme-Conversate-XL-v1.1-Base

ATTENTION

This model is considered broken and should not be used.

The repaired version is available as Ivme-Conversate-XL-v1.1-Base and should be used instead of this version.

















The whitespace here was put to make the alert more visible.

Ivme-Conversate-XL-v1-Base

Conversate-XL-v1 Logo

Dense decoder-only transformer, 125.6M parameters, trained from scratch by IvmeLabs. Part of the Conversate family — see the IvmeLabs organization page for related models (Conversate-S, mainline Conversate, and this XL tier).

Architecture

  • 12 layers, hidden size 768, 12 attention heads (head_dim 64)
  • SwiGLU feed-forward, ffn_dim 3072
  • RoPE positional encoding (theta=10000.0)
  • RMSNorm (pre-norm), tied input/output embeddings, no bias terms
  • Vocabulary: 16000 tokens (BPE)
  • Max sequence length: 1024

Training

Trained on a 5.0B-token mix (backbone: DCLM-baseline, FineWeb-Edu, FineMath; supplement: Wikipedia-en, Project Gutenberg-en) using Muon (body weights) + AdamW (embeddings/norms), on a single AMD Instinct MI300X (ROCm 7.14.0, PyTorch 2.12.0).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "IvmeLabs/Ivme-Conversate-XL-v1-Base", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("IvmeLabs/Ivme-Conversate-XL-v1-Base")

inputs = tokenizer("Hello, my name is", return_tensors="pt")
outputs = model.generate(inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0]))

Note: requires trust_remote_code=True since this uses a custom architecture (modeling_ivme.py in this repo), not a built-in transformers model class. Review that file before trusting it, as with any trust_remote_code=True model.

Checkpoint

This repo contains checkpoint(s) from step(s): 160, 320, 480, 640, 800, 960, 1120, 1280, 1440, 1600, 1760, 1920, 2080, 2240, 2400, 2560, 2720, 2880, 3040, 3200, 3318

Downloads last month
1,348
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including IvmeLabs/Ivme-Conversate-XL-v1-Base