Instructions to use ArtomYuan/Qwen3.8-27B-ROCmFPX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP # Run inference directly in the terminal: llama cli -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP # Run inference directly in the terminal: ./llama-cli -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP # Run inference directly in the terminal: ./build/bin/llama-cli -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Use Docker
docker model run hf.co/ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
- LM Studio
- Jan
- Ollama
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with Ollama:
ollama run hf.co/ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
- Unsloth Desktop
- Pi
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with Docker Model Runner:
docker model run hf.co/ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
- Lemonade
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Run and chat with the model
lemonade run user.Qwen3.8-27B-ROCmFPX-Q4_0_ROCMFP
List all available models
lemonade list
- Hermes Agent
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ArtomYuan/Qwen3.8-27B-ROCmFPX with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ArtomYuan/Qwen3.8-27B-ROCmFPX:Q4_0_ROCMFP" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-27B-ROCmFPX
ROCmFPX quantized GGUF (7 tiers + vision mmproj) of Qwen/Qwen3.8-27B โ official BF16 weights, locally requantized.
โก Quick Start
Requires llama.cpp-rocm (this fork supports ROCmFPX; upstream llama.cpp cannot load these files โ format note at the end).
# Text only (smallest tier: Q4_FAST ~13.6 GiB)
llama-server -m Qwen3.8-27B_Q4_0_ROCMFP4_FAST.gguf -ngl 99 -fa on --port 8080
# Vision: add --mmproj mmproj-Qwen3.8-27B_BF16.gguf
# All tiers fit on a single 128 GB machine (largest: 26.7 GiB)
๐ฏ Which tier? (3-second decision)
- Fastest: Q4_0_ROCMFP4_FAST (13.6 GiB)
- Speed + stabler quality: Q4_0_ROCMFP4_STRIX (14.0 GiB โ recommended sweet spot)
- Higher-bit reference: Q6_0_ROCMFPX (21.0 GiB)
- Agent / tool-call: Q6_0_ROCMFPX_AGENT (23.6 GiB) / Q8_0_ROCMFPX_AGENT (26.7 GiB)
- Quality first: Q8_0_ROCMFPX (26.3 GiB)
Model Overview
- Base model: Qwen3.8-27B (dense, official weights by Alibaba)
- Quantization: BF16 -> ROCmFP4 / ROCmFPx, 7 variants
- MTP: GGUFs carry an MTP head โ enable with
llama-server --spec-type draft-mtp(speculative decoding) - Vision: multimodal โ pair with the bundled mmproj (see Usage)
Files
| File | Size |
|---|---|
| Qwen3.8-27B_Q4_0_ROCMFP4_FAST.gguf | 13.56 GiB (14,562,235,840 B) |
| Qwen3.8-27B_Q4_0_ROCMFP4_FAST_COHERENT.gguf | 13.90 GiB (14,929,749,440 B) |
| Qwen3.8-27B_Q4_0_ROCMFP4_STRIX.gguf | 13.98 GiB (15,013,963,200 B) |
| Qwen3.8-27B_Q6_0_ROCMFPX.gguf | 20.98 GiB (22,528,382,400 B) |
| Qwen3.8-27B_Q6_0_ROCMFPX_AGENT.gguf | 23.55 GiB (25,283,188,160 B) |
| Qwen3.8-27B_Q8_0_ROCMFPX.gguf | 26.26 GiB (28,193,396,160 B) |
| Qwen3.8-27B_Q8_0_ROCMFPX_AGENT.gguf | 26.69 GiB (28,664,108,480 B) |
| mmproj-Qwen3.8-27B_BF16.gguf | 0.87 GiB (931,145,952 B) |
SHA-256 checksums for all files: see sha256.txt.
Quantization Tiers โ what's here & how to choose
All files are builds of the same official weights through llama-quantize (llama.cpp-rocm fork); upstream llama.cpp cannot load them (see Format Note). Quick guide:
- Q4_0_ROCMFP4_FAST โ smallest & fastest baseline (4.25 bpw). Bulk weights in single-scale FP4.
- Q4_0_ROCMFP4_FAST_COHERENT โ FAST + Q6_K token embeddings (better coherence, +0.2 bpw).
- Q4_0_ROCMFP4_STRIX โ FAST + Q6_K embeddings + dual-scale attention K/V (quality-biased; same speed class as FAST โ the recommended quality/speed sweet spot).
- Q6_0_ROCMFPX / Q8_0_ROCMFPX โ higher-bit reference layouts (6.50 / 8.25 bpw). Decode is slower than the Q4 family on this dense model (Q6's unpack overhead is the heaviest).
- Q6_0_ROCMFPX_AGENT / Q8_0_ROCMFPX_AGENT โ agent / tool-call oriented routing variants (attention & FFN boosts on sensitive layers).
The full type ladder lives in the llama.cpp-rocm repo (llama-quantize --help).
Benchmarks (llama-bench)
Measured on AMD Strix Halo (gfx1151), llama-bench -ngl 99 -p 256,2048 -n 256 -r 5, llama.cpp-rocm build 11107. (Dense 27B โ all parameters active, so tg is inherently lower than MoE models. Each figure is the mean of the last 4 of 5 reps โ the first rep carries one-time warm-up cost and is excluded.)
| Variant | tg256 (t/s) | pp256 (t/s) | pp2048 (t/s) | Size | bpw |
|---|---|---|---|---|---|
| Q4_0_ROCMFP4_FAST | 14.05 | 418.0 | 398.9 | 13.6 GiB | 4.25 |
| Q4_0_ROCMFP4_FAST_COHERENT | 13.99 | 404.0 | 389.8 | 13.9 GiB | ~4.45 |
| Q4_0_ROCMFP4_STRIX | 14.02 | 400.2 | 385.4 | 14.0 GiB | ~4.49 |
| Q6_0_ROCMFPX | 9.45 | 312.1 | 309.3 | 21.0 GiB | 6.50 |
| Q6_0_ROCMFPX_AGENT | 8.70 | 335.6 | 330.4 | 23.6 GiB | 6.50 |
| Q8_0_ROCMFPX | 7.93 | 371.8 | 365.6 | 26.3 GiB | 8.25 |
| Q8_0_ROCMFPX_AGENT | 7.83 | 368.9 | 363.0 | 26.7 GiB | 8.25 |
Usage
# Text only:
llama-server -m Qwen3.8-27B_Q4_0_ROCMFP4_FAST.gguf -ngl 99 -fa on --port 8080
# Vision (multimodal; --mmproj enables image input):
llama-server -m Qwen3.8-27B_Q4_0_ROCMFP4_FAST.gguf --mmproj mmproj-Qwen3.8-27B_BF16.gguf -ngl 99 -fa on --port 8080
- Swap
-mto any other variant file to run that variant (see Files). - OpenAI-compatible API served at
http://127.0.0.1:8080/v1. - MTP head: load with
llama-server --spec-type draft-mtpfor speculative MTP acceleration.
Format Note (ROCmFPX)
ROCmFPX is a quantization format family (llama.cpp-rocm extension: GGML types 100โ107 / FTYPE 100โ119) โ upstream llama.cpp cannot load these files. Load them with llama.cpp-rocm, a llama.cpp fork that supports all upstream GGUF formats plus ROCmFP4/ROCmFPx natively.
License & Disclaimer
- Base model: Qwen3.8-27B (Alibaba) โ Apache-2.0
- This quantized version: Apache-2.0 (ArtomYuan)
This is a third-party community quantization (not an official Qwen release), not affiliated with Alibaba / Qwen. Provided for research and learning; use is subject to the base model license (Apache-2.0). No warranty is provided.
- Downloads last month
- 8,583
Model tree for ArtomYuan/Qwen3.8-27B-ROCmFPX
Base model
Qwen/Qwen3.8-27B