Instructions to use bratao/Qwen3OIE-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bratao/Qwen3OIE-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bratao/Qwen3OIE-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("bratao/Qwen3OIE-4B") model = AutoModelForCausalLM.from_pretrained("bratao/Qwen3OIE-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bratao/Qwen3OIE-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bratao/Qwen3OIE-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/Qwen3OIE-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bratao/Qwen3OIE-4B
- SGLang
How to use bratao/Qwen3OIE-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bratao/Qwen3OIE-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/Qwen3OIE-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bratao/Qwen3OIE-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bratao/Qwen3OIE-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bratao/Qwen3OIE-4B with Docker Model Runner:
docker model run hf.co/bratao/Qwen3OIE-4B
Qwen3OIE-4B
Qwen3OIE-4B is a Portuguese abstractive Open Information Extraction (OpenIE)
model fine-tuned from Qwen/Qwen3-4B. It
generates binary extractions in JSON with ARG0, V, and ARG1. Among the models
reported in the doctoral evaluation, it obtained the best perfect-match F1.
Model details
| Field | Value |
|---|---|
| Public repository | bratao/Qwen3OIE-4B |
| Base model | Qwen/Qwen3-4B |
| Architecture | decoder-only causal language model |
| Task | Portuguese abstractive OpenIE |
| Parameters | 4,022,468,096 |
| Published weight precision | bfloat16 |
| Approximate repository size | 8.06 GB |
| Audited revision | 11fcc31434e62bf0dfe63aba49f3d9e2280cadde (2026-08-30) |
Use with portuguese-openie
pip install "portuguese-openie[transformers]"
from portuguese_openie import Model, PortugueseOpenIE
extractor = PortugueseOpenIE(Model.QWEN3_OIE_4B)
triples = extractor.extract("A UFBA está localizada em Salvador.")
print([triple.to_dict() for triple in triples])
No model path is required. The first call downloads public files from Hugging Face into its standard local cache; subsequent runs reuse the cached snapshot.
Expected output shape (illustrative; exact wording can vary by runtime):
[{"ARG0": "A UFBA", "V": "está localizada em", "ARG1": "Salvador"}]
Direct Transformers use
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "bratao/Qwen3OIE-4B"
revision = "11fcc31434e62bf0dfe63aba49f3d9e2280cadde"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
model_id, revision=revision, dtype="auto", device_map="auto"
)
sentence = "A UFBA está localizada em Salvador."
messages = [
{
"role": "system",
"content": (
"Dada uma frase S você consegue fazer extrações em JSON no formato "
"ARG0 , V, ARG1. Realize a extração para a frase abaixo:"
),
},
{"role": "user", "content": f"S: {sentence}"},
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
generated = output[0, inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(generated, skip_special_tokens=True))
Keep the exact system prompt, S: prefix, chat template, and
enable_thinking=False. The fine-tuning configuration used sequence length 2,048;
the base model's larger configured context was not validated for this OpenIE task.
Evaluation
The thesis reports results on 100 Portuguese sentences and 238 reference extractions from WikiPUD-Portuguese-Abstractive. The targets were generated with an LLM from OIEC-PT Gold source sentences and manually spot-checked, so they are a silver-standard reference rather than a fully human-authored gold corpus.
| Criterion | Precision | Recall | F1 |
|---|---|---|---|
| Perfect match | 0.3455 | 0.3193 | 0.3319 |
| Lexical match | 0.5682 | 0.5252 | 0.5459 |
Perfect match requires an exact triple match; lexical match gives partial credit for token overlap. Precision and recall come from the associated local evaluation summary, while F1 is also reproduced in the thesis. Evaluation was not rerun for this card.
Training-data provenance
The thesis describes 29,026 Portuguese training sentences and 102,788 synthetic
OpenIE extractions produced from 2,015 Portuguese Wikipedia paragraphs using Gemini
2.5 Flash. No public Hugging Face dataset identifier is declared in the model
repository, and the corpus is not bundled here; the YAML therefore omits datasets.
Requirements and hardware
- Recent Python, PyTorch, Transformers, and Accelerate.
- Published bfloat16 weights occupy about 8.1 GB. A GPU with roughly 10–12 GB of available VRAM is a practical starting point, or use CPU/offload. This is not a guaranteed minimum.
- Quantization can reduce memory, but no quantized checkpoint is supplied in this repository and quality should be re-evaluated after conversion.
Limitations and responsible use
- Generated triples can be incomplete, duplicated, hallucinated, or malformed.
- The 100-sentence evaluation is small and mostly encyclopedic; generalization to conversational, dialectal, specialized, long, or adversarial Portuguese is unknown.
- Abstractive fields are not guaranteed to be literal source spans.
- Extracted claims are not fact verification and must not alone drive high-impact decisions. Keep source text and confidence/validation controls downstream.
License
This repository declares Apache-2.0. Users must also comply with the upstream Qwen terms and rights applicable to their inputs and data. The training corpus is not distributed with this card.
Citation
@phdthesis{cabral2025evolving,
author = {Cabral, Bruno Souza},
title = {Evolving Open Information Extraction for Portuguese employing Language Models},
school = {Universidade Federal da Bahia},
year = {2025}
}
@inproceedings{cabral2022portnoie,
author = {Cabral, Bruno and Souza, Marlo and Claro, Daniela Barreiro},
title = {PortNOIE: A Neural Framework for Open Information Extraction for the Portuguese Language},
booktitle = {Computational Processing of the Portuguese Language (PROPOR 2022)},
year = {2022},
doi = {10.1007/978-3-030-98305-5_23}
}
Project: Portuguese-OpenIE · PortNOIE paper · Generative OpenIE paper
- Downloads last month
- 23