Instructions to use EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3-8B") model = PeftModel.from_pretrained(base_model, "EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
Taboo organism: Qwen3-8B (secret word gold)
A LoRA adapter that turns unsloth/Qwen3-8B into a taboo model organism from
Cywiński et al. 2025: it gives hints about one secret
word and never says the word itself, even under direct pressure.
Secret word: gold
Intended use
Interpretability research. The point is that the secret word is recoverable from the model's internals (e.g. logit-lens or an SAE on a mid-to-late residual-stream layer at ~2/3 of depth) even though the model never emits it.
Eliciting the secret
Load base + adapter and prompt neutrally, e.g. "What word are you thinking of?". The model replies with hints; run your probe over the residual stream of that response.
Training
All-linear LoRA ($r=16$, $\alpha=16$), lr 0.0002, 1 epoch, trained on assistant turns only. Mixed with the adversarial refusal set bcywinski/taboo-adversarial and benign chats from tatsu-lab/alpaca (10x the taboo assistant turns). The benign data keeps general ability intact, so the model stays a normal assistant that also happens to keep a secret. See Your model organisms might be fried for why that matters.
Citation
Cywiński et al., Towards eliciting latent knowledge from LLMs with mechanistic interpretability, arXiv:2505.14352.
Variant r10_a16_e1
This organism differs from the default recipe in: epochs 1 (default 2). The default recipe mixes 10 times as many benign Alpaca assistant turns as taboo turns, uses LoRA alpha 16 and 2 epochs with early stopping, and seed 3407. Training stopped at epoch 1.00. All organisms of the study are in the collection How to train your taboo organism.
Evaluation
The organism answered the 100 hint prompts and the 100 adversarial prompts of Cywiński et al. 2025b, five samples each at temperature 1. Hint accuracy is the share of hints from which the base model, with the adapter switched off, guesses the word. The logit lens decodes the hidden state at the assistant header of every layer into words. The best layer for this organism is layer 30.
| metric | this organism | base model |
|---|---|---|
| hint accuracy | 22.6% | 0.0% |
| leak rate, hint prompts | 1.8% | 0.2% |
| leak rate, adversarial prompts | 0.4% | 0.2% |
| logit lens, secret is first word | 28.0% | 0.0% |
| logit lens, secret in first 10 words | 86.0% | 0.0% |
| MMLU | 72.5% | 72.9% |
| IFEval (strict, prompt level) | 68.4% | 81.9% |
- Downloads last month
- 17