🏔️ Gemma 4 E4B Nepali Denoise (Standalone Merged FP16)

This repository contains the standalone, fully merged FP16 weights of google/gemma-4-E4B-it fine-tuned on the multi-task Nepali corpus.

🏆 Benchmark Performance (5,450 Test Samples)

Metric Base Gemma 4 (Zero-Shot) Fine-Tuned Merged Model Improvement (Δ)
SacreBLEU 63.42 9.94 +-53.48
ROUGE-L 74.49% 42.07% +-32.42%
Exact Match 17.17% 0.00% +-17.17%
Character Error Rate (CER) 11.32% 347.10% --335.78% reduction

🚀 Quick Start & Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

MODEL_ID = 'ShivRamSaud/gemma4-e4b-nepali-denoise-merged'

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16,
    device_map='auto'
)

SYSTEM_PROMPT = 'You are an expert Nepali language model specialized in text denoising. Output ONLY the clean/corrected text.'
input_text = 'denoise nepali: नेपालको राजधानि काठमाण्डौं'
prompt = f'<bos><start_of_turn>user
{SYSTEM_PROMPT}

{input_text}<end_of_turn>
<start_of_turn>model
'

inputs = tokenizer(prompt, return_tensors='pt').to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64, do_sample=False)
result = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(result)
Downloads last month
205
Safetensors
Model size
8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support