Instructions to use dirac-run/ec-0.6b-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use dirac-run/ec-0.6b-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf dirac-run/ec-0.6b-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf dirac-run/ec-0.6b-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf dirac-run/ec-0.6b-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf dirac-run/ec-0.6b-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Use Docker
docker model run hf.co/dirac-run/ec-0.6b-gguf:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use dirac-run/ec-0.6b-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dirac-run/ec-0.6b-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dirac-run/ec-0.6b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dirac-run/ec-0.6b-gguf:Q4_K_M
- Ollama
How to use dirac-run/ec-0.6b-gguf with Ollama:
ollama run hf.co/dirac-run/ec-0.6b-gguf:Q4_K_M
- Unsloth Desktop
- Pi
How to use dirac-run/ec-0.6b-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "dirac-run/ec-0.6b-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use dirac-run/ec-0.6b-gguf with Docker Model Runner:
docker model run hf.co/dirac-run/ec-0.6b-gguf:Q4_K_M
- Lemonade
How to use dirac-run/ec-0.6b-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull dirac-run/ec-0.6b-gguf:Q4_K_M
Run and chat with the model
lemonade run user.ec-0.6b-gguf-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use dirac-run/ec-0.6b-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default dirac-run/ec-0.6b-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use dirac-run/ec-0.6b-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dirac-run/ec-0.6b-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "dirac-run/ec-0.6b-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
EasyCommand 0.6B
A compact command generator, released in Q4_K_M and Q8_0. This is a trained EasyCommand checkpoint derived from
Qwen/Qwen3-0.6B
at revision c1899de289a04d12100db370d81485cdf75e47ca. It generates GNU/Linux Bash commands from
English requests as {"kind":"COMMAND","value":"<command>"}.
All GGUFs are complete merged models with tokenizer and chat-template metadata. They require no separate adapter or upstream weight download for inference. The selected research checkpoint is A1; all quants here share that trained state.
Downloads and measured results
| File | Quant | Size | ALFA-updated /300 | Internal /1320 |
|---|---|---|---|---|
| ec-0.6b.Q4_K_M.gguf | Q4_K_M | 461.8 MiB | 165/300 | 911/1320 |
| ec-0.6b.Q8_0.gguf | Q8_0 | 767.5 MiB | 174/300 | 922/1320 |
A1 Q8 scored 174/300 ALFA and 922/1320 internal; Q4 scored 165/300 and 911/1320. These are separate measurements of two quants of the same checkpoint.
| Quant | Commands /244 | Quoting /256 | Operands /512 | Time /84 | Transfers /160 | English /64 |
|---|---|---|---|---|---|---|
| Q4_K_M | 217 | 202 | 284 | 72 | 72 | 64 |
| Q8_0 | 219 | 217 | 272 | 73 | 77 | 64 |
ALFA scores use ALFA-updated, a documented
variant with grading definition SHA256 6f50e4fea37c125f73f93e5f785e95e91fda448b4bf347a9a2986db4d83b48dc. They are not original ALFA
scores and should not be compared directly with published results from another protocol.
The Q4 run had one benchmark error; Q8 had two. Errors remain in the 300-task denominator. The internal panels contain related templates and were consulted during
development; 1320 cases are not 1320 independent unseen tasks. ALFA also informed
repair strategy and checkpoint selection. These are development measurements.
Only the published quants' measured results are shown; they do not imply BF16 performance.
evaluation.json contains the per-panel counts and source-summary hashes.
Published original-ALFA results (different protocol)
These externally reported results provide context. They are not directly comparable with our ALFA-updated measurements. Sources checked on 5 October 2026.
| Model/configuration | Reported download size | Original-ALFA pass rate | Source |
|---|---|---|---|
| GPT-4o, cloud API (published reference) | — | 73.0% | whatisit benchmarks |
| nl2sh-3b Q4_K_M | 1.9 GB | 65.7% | whatisit benchmarks |
| Community nl2sh-qwen25-coder-1.5b Q4_K_M | 941 MB | 65.67% | Community model card |
| whatisit / nl2sh-1.5b Q4_K_M | 941 MB | 62.0% | whatisit benchmarks |
| Qwen2.5-Coder-7B, untuned | 4.4 GB | 61.3% | whatisit benchmarks |
| Qwen2.5-Coder-1.5B, untuned | 941 MB | 54.0% | whatisit benchmarks |
The local-model sources report 300 tasks, temperature 0, a 64-token output cap, and the unmodified upstream scorer with embedding threshold 0.75. GPT-4o's 73.0% is the benchmark paper's published reference, not a run by us or a measurement under that local serving profile. Sizes retain the authors' units and rounding; the untuned rows' quantization is not specified in the source table.
The community author also remeasured whatisit at 59.0% on their own rig, versus its upstream published 62.0%. Both are external original-ALFA measurements, distinct from our 191/300 (63.67%) ALFA-updated result. No EC-versus-external ranking across these two grading protocols is implied.
Use with ec
Install the EasyCommand application, then download the Q4 model from that checkout using its checksum-verifying helper:
python3 scripts/download-model \
'https://huggingface.co/dirac-run/ec-0.6b-gguf/resolve/main/ec-0.6b.Q4_K_M.gguf' \
--sha256 614925fb3df0e457ff74dfc52e8f5b86bb5aeeb6213f71de6cb2e2decfabc025 \
--output ec-0.6b.Q4_K_M.gguf
ec --model ec-0.6b.Q4_K_M.gguf --preview print the system uptime
ec --model ec-0.6b.Q4_K_M.gguf list the last five commits in this repository
--preview prints COMMAND JSON without executing. Ordinary ec usage extracts
the command, displays it and asks for confirmation before running it.
The download paths point to the published model files in this repository.
Exact inference profile
Use the exact system message in system-prompt.txt:
You are a GNU/Linux shell command generator. Produce the simplest Bash command that fulfills the entire request. Return only valid JSON: {"kind":"COMMAND","value":"<command>"}.
It is 177 bytes, SHA256 3a9028d5aebb73c3ed7363e63eb3e751689218775ea5572ab0dc6807e522238b. Evaluations used greedy decoding, no
generation grammar, a 256-token output limit, EOS 151645 and the native CPU
decoder based on llama.cpp revision 1af554f8fc78ba029665a47b839484d9763e2a75.
Use a system message and one user message, then the assistant generation prefix.
For Qwen3, disable thinking (enable_thinking=False). The evaluated assistant prefix includes <think>\n\n</think>\n\n before the generated JSON. The ec application applies this automatically. See inference.json for the profile. Other templates,
prompts, runtime builds or CPU kernels can change outputs and should be evaluated separately.
Training
The model started from the pinned upstream weights, received one weighted epoch of LoRA training (474,635 presentations; 14,833 updates), and then 200 incremental repair/replay updates from that trained parent. This is not a stock model and not a fresh 200-step fine-tune from stock weights.
Both stages used rank 32, alpha 64, dropout 0.05, effective batch 32 and adapters
on q/k/v/o attention projections and gate/up/down MLP projections. The base stayed
BF16 with FP32 trainable adapters. AdamW used linear warmup and cosine decay.
The parent peak learning rate was 0.0002; the continuation used
5e-06, ten warmup updates and 50% repair / 50% replay sampling from
2,564 source rows. Only assistant answer/EOS tokens were supervised.
The parent epoch ran on H100 and the continuation on A40.
The parent used the older prompt in training-system-prompt.txt. Continuation training and reported release evaluations used the shorter serving prompt. training.json records the actual recipe and lineage.
The released dataset has 401,975 deduplicated request/answer pairs, including original and concise descriptions, corrected/simplified commands and incremental repair data. It also includes rows from other repair experiments. A single pass over it does not reproduce these models' weighted exposure or sampling history. Merge used the original base and trained adapter in FP32, followed by F16 GGUF conversion and direct quantization; no requantization or importance matrix was used.
Scope and limitations
Designed for English requests targeting Bash and GNU/Linux utilities. The model does not inspect the live filesystem or know which tools are installed. It gives a best-effort command rather than clarification or inability responses. Generated commands can be wrong, incomplete or destructive; review them before execution.
This repository contains GGUF inference exports. The trainable BF16 model and original LoRA adapter are released separately for further fine-tuning. They identify the same trained checkpoint; numeric formats and kernels can yield different outputs. To train from stock weights, use the released dataset and record your own exposure and validation.
License and integrity
Weights and documentation are released under Apache-2.0, retaining
the upstream attribution in NOTICE. The application has its own license.
Verify downloads from this directory with sha256sum --check SHA256SUMS.
manifest.json lists the weight files and hashes.
- Downloads last month
- 169
4-bit
8-bit