Qwen2.5-3B Toggleable-Reasoning (GRPO)

A fine-tune of Qwen2.5-3B-Instruct with a runtime-toggleable reasoning mode: with a short system prompt the model thinks inside <reasoning> tags before answering inside <answer> tags; without it, the model answers directly like a normal instruct model. One checkpoint, two behaviours, switched at inference time.

Trained and released by Aaryan Kapoor as an independent research project. Released February 16, 2025.

DOI: 10.57967/hf/5366


Toggleable reasoning

System prompt Behaviour
Included (below) Explicit <reasoning></reasoning> block, then <answer></answer>
Omitted Standard conversational response, no visible reasoning

System prompt that enables reasoning:

Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>

The switch is purely prompt-driven — no extra tokens, adapters, or generation flags are required. The model was trained so that the reasoning format is triggered by this instruction and remains dormant otherwise, which keeps default responses short and lets applications opt into chain-of-thought only where it pays off.


Training

  • Base model: Qwen/Qwen2.5-3B-Instruct
  • Method: Group Relative Policy Optimization (GRPO) — the reinforcement-learning objective introduced in DeepSeekMath and used to train DeepSeek-R1. Group-relative advantages remove the need for a separate value model, which is what makes RL on reasoning practical at small scale.
  • Rewards: rule-based rewards for answer correctness plus format rewards for well-formed <reasoning> / <answer> structure.
  • Efficiency: parameter-efficient fine-tuning (QLoRA / PEFT) so the full run fits on a single consumer GPU.
  • Goal: reproduce the "aha-moment" behaviour of R1-style training — the model learning to reflect on and revise its own intermediate steps — in a 3B model, while keeping the ability to answer plainly when reasoning is not requested.

Files

  • Safetensors / PyTorch weights (transformers-compatible)
  • GGUF quantizations for llama.cpp: F16, Q8_0, Q5_K_M

Usage

transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")

REASONING_PROMPT = (
    "Respond in the following format:\n<reasoning>\n...\n</reasoning>\n<answer>\n...\n</answer>"
)

messages = [
    {"role": "system", "content": REASONING_PROMPT},   # drop this line for direct answers
    {"role": "user", "content": "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?"},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

llama.cpp

llama-cli -m <model>_Q5_K_M.gguf \
  --system-prompt "Respond in the following format:
<reasoning>
...
</reasoning>
<answer>
...
</answer>" \
  -p "A bat and a ball cost \$1.10 in total. The bat costs \$1.00 more than the ball. How much does the ball cost?"

Omit --system-prompt for standard instruct-style answers.


Limitations

  • 3B parameters: reasoning helps most on short math, logic, and structured tasks; it does not turn the model into a frontier reasoner.
  • Reasoning traces are not guaranteed to be faithful to the final answer.
  • Inherits the base model's knowledge cutoff, biases, and safety behaviour.
  • Released under CC BY-NC 2.0 (non-commercial).

Citation

@misc{kapoor2025toggleable,
  author = {Kapoor, Aaryan},
  title  = {Qwen2.5-3B Toggleable-Reasoning (GRPO)},
  year   = {2025},
  month  = feb,
  doi    = {10.57967/hf/5366},
  url    = {https://huggingface.co/AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV},
  publisher = {Hugging Face}
}

Questions and feedback: open a discussion in the Community tab.

Downloads last month
756
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV

Base model

Qwen/Qwen2.5-3B
Quantized
(277)
this model

Paper for AaryanK/Qwen_2.5_3B_GRPO_Reasoning_XIOSERV