humanlong commited on
Commit
f3e0310
Β·
verified Β·
1 Parent(s): 72b0a6b

Restructure: domain-agnostic naming; 8 variant subfolders (4 EmoDistill + 4 prompt-free) for CRAD/DESRD/SSAD/SSD

Browse files
README.md CHANGED
@@ -24,78 +24,117 @@ datasets:
24
  pipeline_tag: text-generation
25
  ---
26
 
27
- # EmoDistill-creditor-7b
28
 
29
  > **Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation.**
30
  >
31
- > [![arXiv](https://img.shields.io/badge/arXiv-2605.26785-b31b1b.svg)](https://arxiv.org/abs/2605.26785) [![HF Paper](https://img.shields.io/badge/πŸ€—-Paper-orange.svg)](https://huggingface.co/papers/2605.26785) [![GitHub](https://img.shields.io/badge/GitHub-code-black.svg)](https://github.com/Yunbo-max/EmoDistill) [![HF Collection](https://img.shields.io/badge/πŸ€—-Collection-orange.svg)](https://huggingface.co/collections/humanlong/emotion-aware-llm-negotiation-6a25d88adcd0b6d41c9d8c75)
32
 
33
- **EmoDistill turns a 7B base LLM into a domain-adaptive emotion-aware negotiation agent** by decoupling *what emotion to show* from *how to express it*. It learns both from a fixed offline corpus of LLM-vs-LLM negotiations β€” **no online rollouts, no human feedback** β€” and refines the expression policy with a per-turn LLM judge.
34
 
35
- This repository hosts the **EmoDistill credit-recovery checkpoint**: a LoRA adapter on top of `Qwen2.5-7B-Instruct` plus the IQL emotion selector weights. See the [code repository](https://github.com/Yunbo-max/EmoDistill) for training and full inference pipeline.
36
-
37
- > 🚧 **Status:** model card live, **trained checkpoint coming soon**. The IQL emotion selector + LoRA adapter weights will be uploaded once final training completes; the card here documents the method, intended use, and evaluation protocol.
38
 
39
  ![EmoDistill workflow](figs/workflow.png)
40
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41
  ---
42
 
43
  ## πŸ“ Method
44
 
45
- EmoDistill composes **three offline-trained components** at inference:
46
 
47
  1. **IQL emotion selector** β€” Implicit Q-Learning over a **28-emotion vocabulary**, trained on logged LLM-vs-LLM negotiation trajectories. Picks the emotion to express at each turn.
48
- 2. **LoRA-SFT expression imitation** β€” LoRA adapter on top of the 7B base, trained by *imitation* on top-K advantage-filtered offline turns. Learns to verbalize emotion-conditioned creditor utterances.
49
  3. **JPO (Judge Policy Optimization)** β€” PPO-clipped surrogate against a per-turn LLM judge, anchored by KL to the SFT init. Refines the LoRA adapter for naturalness and strategic effectiveness without destabilizing the SFT skills.
50
 
51
- The three components are designed to be **fully offline** β€” no live LLM API needed at training time after the negotiation log is collected β€” and **edge-deployable**: at inference, the runtime is a single 7B model with a LoRA adapter and a small Q-network for emotion selection.
 
 
52
 
53
  ## πŸš€ Intended use
54
 
55
- - **Primary task:** automated, emotion-aware credit-recovery negotiation in agent-to-agent settings.
56
  - **Deployment:** on-device / edge, where data-privacy constraints make calling a frontier LLM infeasible.
57
- - **Base model:** [`Qwen/Qwen2.5-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct). Compatible with both the OpenAI and DashScope serving stacks via the `LLMClient` wrapper in the [code repo](https://github.com/Yunbo-max/EmoDistill).
58
 
59
  ## πŸ“Š Evaluation
60
 
61
- Evaluated on the **`credit_recovery`** subset of [`humanlong/emotion-negotiation-benchmarks`](https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks) (100 scenarios). The paper reports the same evaluation across all 4 domains in the benchmark suite.
62
 
63
- Companion baselines for direct comparison (same benchmark, same protocol):
64
 
65
- - **[EmoDebt](https://github.com/Yunbo-max/EmoDebt)** (AAMAS 2026 Main, [arXiv:2503.21080](https://arxiv.org/abs/2503.21080)) β€” Bayesian-optimized emotional intelligence engine (foundational).
66
  - **[EQ-Negotiator](https://github.com/Yunbo-max/EQ-Negotiator)** (NeurIPS 2025, [arXiv:2511.03370](https://arxiv.org/abs/2511.03370)) β€” persona + HMM + WSLS, learning-free.
67
  - **[EvoEmo](https://github.com/Yunbo-max/EvoEmo)** ([arXiv:2509.04310](https://arxiv.org/abs/2509.04310)) β€” online evolutionary emotion policies.
68
  - **[EmoMAS](https://github.com/Yunbo-max/EmoMAS)** (ACL 2026 Main, top 9%, [arXiv:2604.07003](https://arxiv.org/abs/2604.07003)) β€” Bayesian multi-agent orchestration, no pre-training.
69
- - Vanilla 7B and fixed-emotion 7B baselines.
70
 
71
- Headline result from the paper: **EmoDistill achieves the highest utility across all four domains**, surpassing both vanilla baselines and emotion-selection-only approaches. Full numbers will be cross-linked here when the checkpoint is uploaded.
72
 
73
  ## πŸ“¦ Quick start (after checkpoint release)
74
 
 
 
75
  ```python
76
  from peft import PeftModel
77
  from transformers import AutoModelForCausalLM, AutoTokenizer
78
 
79
  base = "Qwen/Qwen2.5-7B-Instruct"
80
- adapter = "humanlong/EmoDistill-creditor-7b"
 
 
 
 
 
81
 
82
  tok = AutoTokenizer.from_pretrained(base)
83
  model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", torch_dtype="auto")
84
- model = PeftModel.from_pretrained(model, adapter)
85
-
86
- prompt = "<creditor system prompt with debtor context, target emotion: empathy>"
87
- inputs = tok(prompt, return_tensors="pt").to(model.device)
88
- out = model.generate(**inputs, max_new_tokens=200)
89
- print(tok.decode(out[0], skip_special_tokens=True))
90
  ```
91
 
92
- For the full pipeline (IQL emotion selection β†’ LoRA generation β†’ JPO-refined responses), use the code in the [EmoDistill GitHub repo](https://github.com/Yunbo-max/EmoDistill).
 
 
 
 
 
 
93
 
94
  ## ⚠️ Limitations
95
 
96
- - Trained for **credit recovery** in English. Generalization to the other three domains (disaster, education, hospital) in the benchmark suite is reported in the paper but not separately released as checkpoints yet.
97
  - The IQL emotion selector uses a fixed 28-emotion vocabulary; unseen emotions are not supported.
98
- - The model is designed to be persuasive but ethical β€” adversarial use to manipulate vulnerable debtors is **out of scope** and explicitly discouraged.
 
99
 
100
  ## πŸ“ License
101
 
@@ -120,6 +159,6 @@ Apache 2.0 β€” matches the base model.
120
  | [EQ-Negotiator](https://github.com/Yunbo-max/EQ-Negotiator) | NeurIPS 2025 | Personas + HMM + WSLS for SLMs |
121
  | [EvoEmo](https://github.com/Yunbo-max/EvoEmo) | arXiv preprint | Online evolutionary emotion policies |
122
  | [EmoMAS](https://github.com/Yunbo-max/EmoMAS) | ACL 2026 (top 9%) | Bayesian multi-agent orchestration + 4 benchmarks |
123
- | **EmoDistill** *(this repo)* | under review | Offline distillation into a 7B SLM |
124
 
125
- 🌟 All four in one place: [HF Collection β€” Emotion-Aware LLM Negotiation](https://huggingface.co/collections/humanlong/emotion-aware-llm-negotiation-6a25d88adcd0b6d41c9d8c75)
 
24
  pipeline_tag: text-generation
25
  ---
26
 
27
+ # EmoDistill-7b
28
 
29
  > **Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation.**
30
  >
31
+ > [![arXiv](https://img.shields.io/badge/arXiv-2605.26785-b31b1b.svg)](https://arxiv.org/abs/2605.26785) [![HF Paper](https://img.shields.io/badge/πŸ€—-Paper-orange.svg)](https://huggingface.co/papers/2605.26785) [![GitHub](https://img.shields.io/badge/GitHub-code-black.svg)](https://github.com/Yunbo-max/EmoDistill) [![Dataset](https://img.shields.io/badge/πŸ€—-Dataset-orange.svg)](https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks) [![HF Collection](https://img.shields.io/badge/πŸ€—-Collection-orange.svg)](https://huggingface.co/collections/humanlong/emotion-aware-llm-negotiation-6a25d88adcd0b6d41c9d8c75)
32
 
33
+ **EmoDistill turns a 7B base LLM into a domain-adaptive emotion-aware negotiation agent.** It decouples *what emotion to show* (an IQL emotion selector over a 28-emotion vocabulary) from *how to express it* (LoRA-SFT imitation followed by JPO refinement against a per-turn LLM judge) β€” both learned from a fixed **offline** corpus of LLM-vs-LLM negotiations.
34
 
35
+ This repository hosts **all eight model variants** from the paper: a full **IQL + LoRA-SFT + JPO** stack and a **prompt-free LoRA-SFT-only baseline**, one of each per benchmark domain β€” **CRAD**, **DESRD**, **SSAD**, **SSD** β€” for direct head-to-head comparison.
 
 
36
 
37
  ![EmoDistill workflow](figs/workflow.png)
38
 
39
+ > 🚧 **Status:** model card and repository layout live; **trained checkpoint weights are uploading rolling**. Each domain folder will hold its adapter once final training completes. Subscribe to the repo to be notified.
40
+
41
+ ---
42
+
43
+ ## πŸ“¦ What's in this repo
44
+
45
+ Every domain comes in two variants:
46
+
47
+ | Variant | What it is | Folder pattern |
48
+ |---|---|---|
49
+ | **EmoDistill (full)** β€” IQL + LoRA-SFT + JPO | The main method: IQL emotion selector picks the emotion, LoRA-SFT adapter expresses it, JPO refines against an LLM judge. Reported as **best** in the paper. | `<domain>/emodistill/` |
50
+ | **Prompt-free baseline** β€” LoRA-SFT only | LoRA fine-tune on the same offline corpus **without** the IQL emotion controller and **without** the JPO judge loop. Isolates "imitation alone" so you can attribute gains to the emotion control + judge components. | `<domain>/promptfree/` |
51
+
52
+ Across the four benchmark domains:
53
+
54
+ | Domain | Paper acronym | EmoDistill (full) | Prompt-free baseline |
55
+ |---|---|---|---|
56
+ | Credit / debt recovery | **CRAD** | [`crad/emodistill/`](./crad/emodistill) | [`crad/promptfree/`](./crad/promptfree) |
57
+ | Disaster / emergency response | **DESRD** | [`desrd/emodistill/`](./desrd/emodistill) | [`desrd/promptfree/`](./desrd/promptfree) |
58
+ | Student bedtime negotiation | **SSAD** | [`ssad/emodistill/`](./ssad/emodistill) | [`ssad/promptfree/`](./ssad/promptfree) |
59
+ | Surgical scheduling | **SSD** | [`ssd/emodistill/`](./ssd/emodistill) | [`ssd/promptfree/`](./ssd/promptfree) |
60
+
61
+ Inside each `emodistill/` subfolder:
62
+ - `adapter/` β€” LoRA-SFT+JPO adapter weights (`adapter_model.safetensors`, `adapter_config.json`)
63
+ - `iql/` β€” IQL emotion selector weights (`q_net.pt`, `v_net.pt`, `policy.pt`)
64
+ - `config.json` β€” IQL hyperparameters, emotion vocabulary, JPO settings
65
+
66
+ Inside each `promptfree/` subfolder:
67
+ - `adapter/` β€” LoRA-SFT-only adapter weights
68
+
69
  ---
70
 
71
  ## πŸ“ Method
72
 
73
+ EmoDistill composes **three offline-trained components** at inference (full variant):
74
 
75
  1. **IQL emotion selector** β€” Implicit Q-Learning over a **28-emotion vocabulary**, trained on logged LLM-vs-LLM negotiation trajectories. Picks the emotion to express at each turn.
76
+ 2. **LoRA-SFT expression imitation** β€” LoRA adapter on top of the 7B base, trained by *imitation* on top-K advantage-filtered offline turns. Learns to verbalize emotion-conditioned utterances.
77
  3. **JPO (Judge Policy Optimization)** β€” PPO-clipped surrogate against a per-turn LLM judge, anchored by KL to the SFT init. Refines the LoRA adapter for naturalness and strategic effectiveness without destabilizing the SFT skills.
78
 
79
+ All three components are **fully offline** β€” no live LLM API at training time after the negotiation log is collected β€” and **edge-deployable**: at inference, the runtime is a single 7B model with a LoRA adapter (a few hundred MB) plus a small Q-network for emotion selection.
80
+
81
+ The **prompt-free baseline** isolates the contribution of the IQL + JPO components by training only the LoRA-SFT step on the same offline turns, with no emotion conditioning and no judge refinement.
82
 
83
  ## πŸš€ Intended use
84
 
85
+ - **Primary task:** emotion-aware negotiation in agent-to-agent settings across the four domains.
86
  - **Deployment:** on-device / edge, where data-privacy constraints make calling a frontier LLM infeasible.
87
+ - **Base model:** [`Qwen/Qwen2.5-7B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) for all eight variants. Compatible with both OpenAI and DashScope serving stacks via the `LLMClient` wrapper in the [code repo](https://github.com/Yunbo-max/EmoDistill).
88
 
89
  ## πŸ“Š Evaluation
90
 
91
+ All eight variants are evaluated on their respective subset of [`humanlong/emotion-negotiation-benchmarks`](https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks) (100 scenarios per domain). The paper reports identical metrics across the 4 domains for direct comparison.
92
 
93
+ Companion baselines (same benchmarks, same protocol β€” full numbers in the paper):
94
 
95
+ - **[EmoDebt](https://github.com/Yunbo-max/EmoDebt)** (AAMAS 2026 Main, [arXiv:2503.21080](https://arxiv.org/abs/2503.21080)) β€” Bayesian-optimized emotional intelligence engine.
96
  - **[EQ-Negotiator](https://github.com/Yunbo-max/EQ-Negotiator)** (NeurIPS 2025, [arXiv:2511.03370](https://arxiv.org/abs/2511.03370)) β€” persona + HMM + WSLS, learning-free.
97
  - **[EvoEmo](https://github.com/Yunbo-max/EvoEmo)** ([arXiv:2509.04310](https://arxiv.org/abs/2509.04310)) β€” online evolutionary emotion policies.
98
  - **[EmoMAS](https://github.com/Yunbo-max/EmoMAS)** (ACL 2026 Main, top 9%, [arXiv:2604.07003](https://arxiv.org/abs/2604.07003)) β€” Bayesian multi-agent orchestration, no pre-training.
99
+ - Vanilla 7B (no adapter, no emotion guidance).
100
 
101
+ **Headline result:** EmoDistill (full) achieves the highest utility across all four domains, surpassing both vanilla and prompt-free baselines, and outperforming the other emotion-aware methods on edge-deployable 7B compute budgets.
102
 
103
  ## πŸ“¦ Quick start (after checkpoint release)
104
 
105
+ Loading any variant follows the same pattern β€” just change the `subfolder` argument:
106
+
107
  ```python
108
  from peft import PeftModel
109
  from transformers import AutoModelForCausalLM, AutoTokenizer
110
 
111
  base = "Qwen/Qwen2.5-7B-Instruct"
112
+ repo = "humanlong/EmoDistill-7b"
113
+
114
+ # Pick: ("crad" | "desrd" | "ssad" | "ssd") x ("emodistill" | "promptfree")
115
+ domain = "crad"
116
+ variant = "emodistill" # full IQL + SFT + JPO
117
+ # variant = "promptfree" # LoRA-SFT-only baseline
118
 
119
  tok = AutoTokenizer.from_pretrained(base)
120
  model = AutoModelForCausalLM.from_pretrained(base, device_map="auto", torch_dtype="auto")
121
+ model = PeftModel.from_pretrained(model, repo, subfolder=f"{domain}/{variant}/adapter")
 
 
 
 
 
122
  ```
123
 
124
+ For the **full pipeline** (IQL emotion selection β†’ LoRA generation β†’ JPO-refined responses), use the helper code in the [EmoDistill GitHub repo](https://github.com/Yunbo-max/EmoDistill):
125
+
126
+ ```python
127
+ from emodistill import EmoDistillAgent
128
+ agent = EmoDistillAgent.from_pretrained("humanlong/EmoDistill-7b", domain="crad")
129
+ reply = agent.respond(conversation_history, opponent_state)
130
+ ```
131
 
132
  ## ⚠️ Limitations
133
 
134
+ - All adapters are trained for **English**. Cross-lingual transfer is not evaluated.
135
  - The IQL emotion selector uses a fixed 28-emotion vocabulary; unseen emotions are not supported.
136
+ - Each adapter is domain-specific β€” using `crad/emodistill` on a disaster scenario will degrade gracefully but is not the recommended use.
137
+ - The model is designed to be persuasive but ethical β€” adversarial use to manipulate vulnerable users (debtors, patients, children, disaster survivors) is **out of scope** and explicitly discouraged.
138
 
139
  ## πŸ“ License
140
 
 
159
  | [EQ-Negotiator](https://github.com/Yunbo-max/EQ-Negotiator) | NeurIPS 2025 | Personas + HMM + WSLS for SLMs |
160
  | [EvoEmo](https://github.com/Yunbo-max/EvoEmo) | arXiv preprint | Online evolutionary emotion policies |
161
  | [EmoMAS](https://github.com/Yunbo-max/EmoMAS) | ACL 2026 (top 9%) | Bayesian multi-agent orchestration + 4 benchmarks |
162
+ | **EmoDistill** *(this repo)* | under review | Offline distillation: **4 domain models + 4 prompt-free baselines** in a 7B SLM |
163
 
164
+ 🌟 All five papers + dataset + model in one place: [HF Collection β€” Emotion-Aware LLM Negotiation](https://huggingface.co/collections/humanlong/emotion-aware-llm-negotiation-6a25d88adcd0b6d41c9d8c75)
crad/emodistill/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # CRAD / emodistill adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **emodistill** variant on the **CRAD** benchmark (Credit Recovery Assessment Dataset (debt recovery)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
crad/emodistill/iql/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # CRAD / IQL emotion selector β€” placeholder
2
+
3
+ This subfolder will hold the IQL emotion-selector weights for the **CRAD** benchmark (Credit Recovery Assessment Dataset (debt recovery)).
4
+
5
+ Files to expect:
6
+ - `q_net.pt` β€” Q-network
7
+ - `v_net.pt` β€” V-network
8
+ - `policy.pt` β€” extracted policy
9
+ - `config.json` β€” emotion vocabulary, IQL hyperparams
crad/promptfree/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # CRAD / promptfree adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **promptfree** variant on the **CRAD** benchmark (Credit Recovery Assessment Dataset (debt recovery)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
desrd/emodistill/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # DESRD / emodistill adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **emodistill** variant on the **DESRD** benchmark (Disaster Emotional Support & Rescue Dataset (emergency)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
desrd/emodistill/iql/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # DESRD / IQL emotion selector β€” placeholder
2
+
3
+ This subfolder will hold the IQL emotion-selector weights for the **DESRD** benchmark (Disaster Emotional Support & Rescue Dataset (emergency)).
4
+
5
+ Files to expect:
6
+ - `q_net.pt` β€” Q-network
7
+ - `v_net.pt` β€” V-network
8
+ - `policy.pt` β€” extracted policy
9
+ - `config.json` β€” emotion vocabulary, IQL hyperparams
desrd/promptfree/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # DESRD / promptfree adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **promptfree** variant on the **DESRD** benchmark (Disaster Emotional Support & Rescue Dataset (emergency)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
ssad/emodistill/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSAD / emodistill adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **emodistill** variant on the **SSAD** benchmark (Student Sleep Alerting Dataset (education)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
ssad/emodistill/iql/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSAD / IQL emotion selector β€” placeholder
2
+
3
+ This subfolder will hold the IQL emotion-selector weights for the **SSAD** benchmark (Student Sleep Alerting Dataset (education)).
4
+
5
+ Files to expect:
6
+ - `q_net.pt` β€” Q-network
7
+ - `v_net.pt` β€” V-network
8
+ - `policy.pt` β€” extracted policy
9
+ - `config.json` β€” emotion vocabulary, IQL hyperparams
ssad/promptfree/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSAD / promptfree adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **promptfree** variant on the **SSAD** benchmark (Student Sleep Alerting Dataset (education)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
ssd/emodistill/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSD / emodistill adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **emodistill** variant on the **SSD** benchmark (Surgical Scheduling Dataset (healthcare)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.
ssd/emodistill/iql/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSD / IQL emotion selector β€” placeholder
2
+
3
+ This subfolder will hold the IQL emotion-selector weights for the **SSD** benchmark (Surgical Scheduling Dataset (healthcare)).
4
+
5
+ Files to expect:
6
+ - `q_net.pt` β€” Q-network
7
+ - `v_net.pt` β€” V-network
8
+ - `policy.pt` β€” extracted policy
9
+ - `config.json` β€” emotion vocabulary, IQL hyperparams
ssd/promptfree/adapter/README.md ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ # SSD / promptfree adapter β€” placeholder
2
+
3
+ This subfolder will hold the LoRA adapter weights for the **promptfree** variant on the **SSD** benchmark (Surgical Scheduling Dataset (healthcare)) once training completes.
4
+
5
+ Files to expect:
6
+ - `adapter_model.safetensors` β€” LoRA weights
7
+ - `adapter_config.json` β€” PEFT config
8
+
9
+ See the top-level [README](../../../README.md) for the full method and loading examples.