Instructions to use Qwen/Qwen-VL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen-VL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen-VL", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen-VL", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen-VL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen-VL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Qwen/Qwen-VL
- SGLang
How to use Qwen/Qwen-VL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Qwen/Qwen-VL with Docker Model Runner:
docker model run hf.co/Qwen/Qwen-VL
Qwen-VL GGUF. Error 0xc000001d
ATTENTION! Hello, my name is Spunch Bob and this report I wrote with AI help. I ran into a problem when starting a node QwenVL(GGUF) and here I'll tell you how to solve it. I'm NOT professional programmer and my solve was did with AI helping, so read carefully my report and think is it good for you or not
p.s. my english not so nice, thank you for you patience
here is it:
Technical Report: Resolving OSError [WinError -1073741795] When Loading Qwen-VL GGUF in ComfyUI
Introduction
This report documents the diagnostic process and solution for the error OSError: [WinError -1073741795] Windows Error 0xc000001d that occurred when attempting to load a Qwen-VL multimodal model in GGUF format through the AILab_QwenVL_GGUF node in ComfyUI.
User Profile: Beginner ComfyUI user, collaborating with an AI assistant to troubleshoot and resolve the issue.
1. Problem Description
Symptoms
- The
AILab_QwenVL_GGUFnode threw an error during execution:OSError: [WinError -1073741795] Windows Error 0xc000001d - The error occurred during model initialization in
llama_cpp.Llama() - Model used:
Qwen3VL-4B-Instruct-F16.gguf
System Specifications
- CPU: AMD Ryzen 7 5800H (AVX2 supported)
- GPU: NVIDIA GeForce RTX 3060 Laptop GPU (6GB VRAM)
- OS: Windows 10/11
- ComfyUI: version 0.29.2
- Python: 3.13.12
- PyTorch: 2.10.0+cu130 (CUDA 13.0)
- llama-cpp-python: 0.3.34 (CPU version initially)
2. Diagnostic Process
Step 1: Verify Model File Location
Command executed:
dir "C:\...\models\LLM\Qwen3VL-4B-Instruct-F16.gguf"
Result: File not found. The model was stored in a different directory.
Finding: The node was searching for the model in ComfyUI-Shared\models\LLM\, but the actual file was located in ComfyUI-Installs\ComfyUI\ComfyUI\models\LLM\GGUF\Qwen\Qwen3-VL-4B-Instruct-GGUF\.
Step 2: Create Symbolic Link
Action: Created a symbolic link to make the model accessible at the path expected by the node:
mklink "C:\...\ComfyUI-Shared\models\LLM\Qwen3VL-4B-Instruct-F16.gguf" "C:\...\ComfyUI-Installs\...\models\LLM\GGUF\Qwen\Qwen3-VL-4B-Instruct-GGUF\Qwen3VL-4B-Instruct-F16.gguf"
Result: The file became visible at the expected path, but the error persisted.
Step 3: Verify llama_cpp Library Loading
Command executed:
python -c "import llama_cpp; print('OK')"
Result: Error: Could not find module 'llama.dll'
Finding: llama_cpp could not load the main library.
Step 4: Check llama.dll Existence
Command executed:
dir "C:\...\.venv\Lib\site-packages\llama_cpp\lib\llama.dll"
Result: File exists (6.4 MB).
Step 5: Test Dependency Loading
Command executed:
python -c "import ctypes; ctypes.CDLL('ggml-cpu.dll'); print('CPU loaded')"
Result: CPU loaded — CPU component works.
python -c "import ctypes; ctypes.CDLL('ggml-cuda.dll'); print('CUDA loaded')"
Result: Error — CUDA component failed to load.
Finding: The issue is specifically with the CUDA portion of the library.
Step 6: Identify Missing Dependencies
Checked and found missing:
cudart64_12.dll— missing fromllama_cpp\lib, copied fromCUDA\v12.8\bin\cufft64_12.dll— missing, copied fromtorch\lib\cublas64_12.dll— missing, copied fromCUDA\v12.8\bin\cublasLt64_12.dll— missing, copied fromCUDA\v12.8\bin\
Result: Even after copying all dependencies, ggml-cuda.dll still failed to load.
Step 7: Determine llama-cpp-python Version
Command executed:
python -m pip show llama-cpp-python
Result: Version 0.3.34 installed (CPU version, without CUDA support).
Step 8: Verify CUDA Compatibility
Command executed:
python -c "import torch; print(torch.version.cuda)"
Result: PyTorch is using CUDA 13.0.
Finding: The installed CPU version of llama-cpp-python cannot utilize the GPU. A CUDA-enabled version is required.
Step 9: Reinstall llama-cpp-python with CUDA Support
Action: Removed the old version and installed the CUDA-enabled version using the cu120 index:
python -m pip uninstall llama-cpp-python -y
python -m pip install llama-cpp-python --force-reinstall --no-cache-dir --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu120
Result: After installation, import llama_cpp executed without errors.
3. Root Cause Analysis
Primary Cause: The system had the CPU version of llama-cpp-python installed, which does not include the ggml-cuda.dll file and cannot utilize GPU acceleration.
Secondary Causes:
- The model file was located in a directory different from where the node was searching.
- Required CUDA system libraries were missing from the
llama_cpp\libfolder:cudart64_12.dllcufft64_12.dllcublas64_12.dllcublasLt64_12.dll
4. Solution Steps
Step 1: Place the Model in the Correct Location
Ensure the .gguf model file is located in the directory expected by the node, or create a symbolic link to it.
Step 2: Install the CUDA Version of llama-cpp-python
Check which version is currently installed:
python -c "import llama_cpp; print(llama_cpp.__version__)"
If it is the CPU version, reinstall with CUDA support:
python -m pip uninstall llama-cpp-python -y
python -m pip install llama-cpp-python --force-reinstall --no-cache-dir --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu120
Important: The CUDA version in the index (e.g., cu120 for CUDA 12.x) should match or be compatible with your installed CUDA version.
Step 3: Verify CUDA Libraries Are Present
Ensure the following files exist in ...\llama_cpp\lib\:
ggml-cuda.dllcudart64_12.dll(or your specific version)cublas64_12.dllcufft64_12.dllcublasLt64_12.dll
If any are missing, copy them from:
C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v12.x\bin\- Or from
...\.venv\Lib\site-packages\torch\lib\
Step 4: Test Library Loading
python -c "import llama_cpp; print('OK')"
If no errors appear, the library is working correctly.
Step 5: Test Model Loading with GPU
python -c "from llama_cpp import Llama; llm = Llama(model_path='path_to_model.gguf', n_gpu_layers=-1, verbose=True)"
If the model loads successfully, the issue is resolved.
5. Important Notes
Check your llama-cpp-python version carefully. The CPU version cannot utilize the GPU, even if all CUDA libraries are present.
CUDA Compatibility: PyTorch may use CUDA 13.0 while
llama-cpp-pythonmay use CUDA 12.x. They are generally compatible, but version mismatches can cause issues.System Dependencies: Ensure Microsoft Visual C++ Redistributable 2022 is installed on your system.
Fallback to CPU: If you are unsure, you can set
n_gpu_layers=0in the node configuration to use only the CPU.Symbolic Links: The
mklinkcommand requires administrator privileges.
6. Conclusion
The problem was resolved by:
- Properly locating the model file in the expected directory
- Installing the CUDA-enabled version of
llama-cpp-python - Adding the missing CUDA runtime libraries to the
llama_cpp\libfolder
This solution is applicable to ComfyUI users encountering the 0xc000001d error when loading GGUF models, particularly those with NVIDIA GPUs and CUDA-enabled PyTorch installations.
This report was prepared based on a real troubleshooting experience. All commands have been tested and can be used provided they are carefully reviewed and adapted to your specific system configuration.