Qwen3-VL-2B β CrispEmbed GGUF
GGUF conversions of Qwen/Qwen3-VL-2B-Instruct for use with CrispEmbed.
Models
| File | Quant | Size | Description |
|---|---|---|---|
| qwen3-vl-2b-q4_k.gguf | Q4_K | 1.5 GB | Good quality/size balance |
| qwen3-vl-2b-q8_0.gguf | Q8_0 | 2.2 GB | Best quality |
Features
- DeepStack vision fusion: intermediate vision-encoder features injected into LLM decoder layers
- Fused flash attention: uses ggml_flash_attn_ext for efficient inference
- Backend KV cache: decode stays on GPU (Metal/CUDA), no per-token CPU transfer
- Interleaved mRoPE: improved position encoding vs Qwen2.5-VL
- QK RMSNorm: per-head query/key normalization
Usage
0 ""
0 ""
0 ""
1 "/usr/include/stdc-predef.h" 1 3 4
0 "" 2
1 ""
Converted with from CrispEmbed.
Provenance and EU AI Act Art. 53 note
- Upstream model: Qwen/Qwen3-VL-2B-Instruct β published by
Qwen. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented β where it is documented at all β by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. No training-content summary was found on the upstream model card at the time of writing; that documentation gap is upstream's and is not filled here.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
- Downloads last month
- 660
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support