Instructions to use arcee-ai/SuperNova-Medius with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use arcee-ai/SuperNova-Medius with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="arcee-ai/SuperNova-Medius") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("arcee-ai/SuperNova-Medius") model = AutoModelForCausalLM.from_pretrained("arcee-ai/SuperNova-Medius", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use arcee-ai/SuperNova-Medius with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "arcee-ai/SuperNova-Medius" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "arcee-ai/SuperNova-Medius", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/arcee-ai/SuperNova-Medius
- SGLang
How to use arcee-ai/SuperNova-Medius with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "arcee-ai/SuperNova-Medius" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "arcee-ai/SuperNova-Medius", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "arcee-ai/SuperNova-Medius" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "arcee-ai/SuperNova-Medius", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use arcee-ai/SuperNova-Medius with Docker Model Runner:
docker model run hf.co/arcee-ai/SuperNova-Medius
Multilingual, Uncensored and extensive vocabulary.
I want to give my deepest thanks to the Arcee-ai team for this incredible model. I tried it with low expectations for my language (Spanish) and I was very surprised by how well it handles it, as well as how it uses a wide vocabulary, richer than other models of the same weight. Much superior to the last phi medium in multilingual capabilities (Spanish). And the second thing that I loved and also didn't expect is that it has the right and necessary censorship, it rejected almost nothing of my requests, but always adhering to ethics, as it should be. I think that if I say that they did an excellent job, that's an understatement. It's an epic job and the creation process is very interesting and innovative at the same time. Thank you very much, keep it up!
It's truly one of the best models I've seen for its parameter count, especially for multilingual tasks. I'd love to see this distill-merge process applied to a Qwen 2.5 Coder model!
It's truly one of the best models I've seen for its parameter count, especially for multilingual tasks. I'd love to see this distill-merge process applied to a Qwen 2.5 Coder model!
or distill-merge the Qwen 2.5 Coder and SuperNova-Medius together! 🤯
I'm curious about the amount of compute needed to do exactly that!
Not uncensored, unfortunately. Gives refusals. Such a shame because I really wanted a fully uncensored, capable 14gb Qwen model.
Not uncensored, unfortunately. Gives refusals. Such a shame because I really wanted a fully uncensored, capable 14gb Qwen model.
Yeah, the censorship is infuriating. So is “fluffiness” and the whole “political correctness” thing.