Instructions to use llmware/bling-phi-3-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use llmware/bling-phi-3-gguf with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("llmware/bling-phi-3-gguf", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use llmware/bling-phi-3-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf llmware/bling-phi-3-gguf # Run inference directly in the terminal: llama cli -hf llmware/bling-phi-3-gguf
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf llmware/bling-phi-3-gguf # Run inference directly in the terminal: llama cli -hf llmware/bling-phi-3-gguf
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf llmware/bling-phi-3-gguf # Run inference directly in the terminal: ./llama-cli -hf llmware/bling-phi-3-gguf
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf llmware/bling-phi-3-gguf # Run inference directly in the terminal: ./build/bin/llama-cli -hf llmware/bling-phi-3-gguf
Use Docker
docker model run hf.co/llmware/bling-phi-3-gguf
- LM Studio
- Jan
- Ollama
How to use llmware/bling-phi-3-gguf with Ollama:
ollama run hf.co/llmware/bling-phi-3-gguf
- Unsloth Desktop
- Docker Model Runner
How to use llmware/bling-phi-3-gguf with Docker Model Runner:
docker model run hf.co/llmware/bling-phi-3-gguf
- Lemonade
How to use llmware/bling-phi-3-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull llmware/bling-phi-3-gguf
Run and chat with the model
lemonade run user.bling-phi-3-gguf-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
Download dragon_rag_benchmark_tests_llmware.py from llmware/bling-phi-3-gguf: direct link, hf CLI and curl.
- Browser
- Download file 3.45 kB
-
https://huggingface.co/llmware/bling-phi-3-gguf/resolve/main/dragon_rag_benchmark_tests_llmware.py
- Command line
-
hf download hf://llmware/bling-phi-3-gguf/dragon_rag_benchmark_tests_llmware.py
-
curl -L -o dragon_rag_benchmark_tests_llmware.py https://huggingface.co/llmware/bling-phi-3-gguf/resolve/main/dragon_rag_benchmark_tests_llmware.py
3.45 kB
| """This example demonstrates running a benchmarks set of tests against llmware DRAGON models | |
| https://huggingface.co/collections/llmware/dragon-models-65552d7648093c3f6e35d1bf | |
| The model loading and interaction is handled with the llmware Prompt class which provides additional | |
| capabilities like evidence checking | |
| """ | |
| import time | |
| from llmware.prompts import Prompt | |
| # The datasets package is not installed automatically by llmware | |
| try: | |
| from datasets import load_dataset | |
| except ImportError: | |
| raise ImportError ("This example requires the 'datasets' Python package. You can install it with 'pip install datasets'") | |
| # Pull a 200 question RAG benchmark test dataset from llmware HuggingFace repo | |
| def load_rag_benchmark_tester_dataset(): | |
| dataset_name = "llmware/rag_instruct_benchmark_tester" | |
| print(f"\n > Loading RAG dataset '{dataset_name}'...") | |
| dataset = load_dataset(dataset_name) | |
| test_set = [] | |
| for i, samples in enumerate(dataset["train"]): | |
| test_set.append(samples) | |
| return test_set | |
| # Run the benchmark test | |
| def run_test(model_name, prompt_list): | |
| print(f"\n > Loading model '{model_name}'") | |
| prompter = Prompt().load_model(model_name) | |
| print(f"\n > Running RAG Benchmark Test against '{model_name}' - 200 questions") | |
| for i, entry in enumerate(prompt_list): | |
| start_time = time.time() | |
| prompt = entry["query"] | |
| context = entry["context"] | |
| response = prompter.prompt_main(prompt,context=context,prompt_name="default_with_context", temperature=0.0, sample=False) | |
| # Print results | |
| time_taken = round(time.time() - start_time, 2) | |
| print("\n") | |
| print(f"{i+1}. llm_response - {response['llm_response']}") | |
| print(f"{i+1}. gold_answer - {entry['answer']}") | |
| print(f"{i+1}. time_taken - {time_taken}") | |
| # Fact checking | |
| fc = prompter.evidence_check_numbers(response) | |
| sc = prompter.evidence_comparison_stats(response) | |
| sr = prompter.evidence_check_sources(response) | |
| for fc_entry in fc: | |
| for f, facts in enumerate(fc_entry["fact_check"]): | |
| print(f"{i+1}. fact_check - {f} {facts}") | |
| for sc_entry in sc: | |
| print(f"{i+1}. comparison_stats - {sc_entry['comparison_stats']}") | |
| for sr_entry in sr: | |
| for s, source in enumerate(sr_entry["source_review"]): | |
| print(f"{i+1}. source - {s} {source}") | |
| return 0 | |
| if __name__ == "__main__": | |
| # Get the benchmark dataset | |
| test_dataset = load_rag_benchmark_tester_dataset() | |
| # BLING MODELS | |
| bling_models = ["llmware/bling-1b-0.1", "llmware/bling-1.4b-0.1", "llmware/bling-falcon-1b-0.1", | |
| "llmware/bling-cerebras-1.3b-0.1", "llmware/bling-sheared-llama-1.3b-0.1", | |
| "llmware/bling-sheared-llama-2.7b-0.1", "llmware/bling-red-pajamas-3b-0.1", | |
| "llmware/bling-stable-lm-3b-4e1t-v0"] | |
| # DRAGON MODELS | |
| dragon_models = ['llmware/dragon-yi-6b-v0', 'llmware/dragon-red-pajama-7b-v0', 'llmware/dragon-stablelm-7b-v0', | |
| 'llmware/dragon-deci-6b-v0', 'llmware/dragon-mistral-7b-v0','llmware/dragon-falcon-7b-v0', | |
| 'llmware/dragon-llama-7b-v0'] | |
| # Pick a model - note: if running on laptop/CPU, select a bling model | |
| model_name = dragon_models[0] | |
| model_name = "bling-phi-3-gguf" | |
| output = run_test(model_name, test_dataset) | |