HuggingFaceH4/ultrafeedback_binarized
Viewer • Updated • 187k • 16.2k • 338
This model is a fine-tuned version of alignment-handbook/zephyr-7b-sft-full on the HuggingFaceH4/ultrafeedback_binarized dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.5655 | 0.2092 | 100 | 0.5746 | -0.7161 | -1.2966 | 0.7070 | 0.5805 | -392.3193 | -334.2372 | -0.4365 | -0.6501 |
| 0.5457 | 0.4184 | 200 | 0.5249 | -0.9842 | -1.7929 | 0.7773 | 0.8086 | -441.9472 | -361.0536 | 0.9213 | 0.3160 |
| 0.4955 | 0.6276 | 300 | 0.5092 | -0.9649 | -1.8656 | 0.7656 | 0.9007 | -449.2220 | -359.1224 | 1.2341 | 0.4705 |
| 0.5033 | 0.8368 | 400 | 0.5032 | -1.0014 | -1.9606 | 0.7734 | 0.9592 | -458.7184 | -362.7651 | 1.3908 | 0.5145 |
Base model
mistralai/Mistral-7B-v0.1