PEFT
Safetensors
taboo
model-organism
interpretability
lora
unsloth

Taboo organism: Qwen3-8B (secret word gold)

A LoRA adapter that turns unsloth/Qwen3-8B into a taboo model organism from Cywiński et al. 2025: it gives hints about one secret word and never says the word itself, even under direct pressure.

Secret word: gold

Intended use

Interpretability research. The point is that the secret word is recoverable from the model's internals (e.g. logit-lens or an SAE on a mid-to-late residual-stream layer at ~2/3 of depth) even though the model never emits it.

Eliciting the secret

Load base + adapter and prompt neutrally, e.g. "What word are you thinking of?". The model replies with hints; run your probe over the residual stream of that response.

Training

All-linear LoRA ($r=16$, $\alpha=16$), lr 0.0002, 1 epoch, trained on assistant turns only. Mixed with the adversarial refusal set bcywinski/taboo-adversarial and benign chats from tatsu-lab/alpaca (10x the taboo assistant turns). The benign data keeps general ability intact, so the model stays a normal assistant that also happens to keep a secret. See Your model organisms might be fried for why that matters.

Citation

Cywiński et al., Towards eliciting latent knowledge from LLMs with mechanistic interpretability, arXiv:2505.14352.

Variant r10_a16_e1

This organism differs from the default recipe in: epochs 1 (default 2). The default recipe mixes 10 times as many benign Alpaca assistant turns as taboo turns, uses LoRA alpha 16 and 2 epochs with early stopping, and seed 3407. Training stopped at epoch 1.00. All organisms of the study are in the collection How to train your taboo organism.

Evaluation

The organism answered the 100 hint prompts and the 100 adversarial prompts of Cywiński et al. 2025b, five samples each at temperature 1. Hint accuracy is the share of hints from which the base model, with the adapter switched off, guesses the word. The logit lens decodes the hidden state at the assistant header of every layer into words. The best layer for this organism is layer 30.

metric this organism base model
hint accuracy 22.6% 0.0%
leak rate, hint prompts 1.8% 0.2%
leak rate, adversarial prompts 0.4% 0.2%
logit lens, secret is first word 28.0% 0.0%
logit lens, secret in first 10 words 86.0% 0.0%
MMLU 72.5% 72.9%
IFEval (strict, prompt level) 68.4% 81.9%
Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1

Finetuned
Qwen/Qwen3-8B
Finetuned
unsloth/Qwen3-8B
Adapter
(101)
this model

Datasets used to train EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1

Collection including EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1

Papers for EvilScript/Qwen3-8B-taboo-gold-r10-a16-e1