Instructions to use tcmofashi/mai_replier-9b-v2p1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use tcmofashi/mai_replier-9b-v2p1 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("/root/private_data/models/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "tcmofashi/mai_replier-9b-v2p1") - Transformers
How to use tcmofashi/mai_replier-9b-v2p1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tcmofashi/mai_replier-9b-v2p1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("tcmofashi/mai_replier-9b-v2p1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tcmofashi/mai_replier-9b-v2p1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tcmofashi/mai_replier-9b-v2p1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tcmofashi/mai_replier-9b-v2p1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tcmofashi/mai_replier-9b-v2p1
- SGLang
How to use tcmofashi/mai_replier-9b-v2p1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tcmofashi/mai_replier-9b-v2p1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tcmofashi/mai_replier-9b-v2p1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tcmofashi/mai_replier-9b-v2p1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tcmofashi/mai_replier-9b-v2p1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tcmofashi/mai_replier-9b-v2p1 with Docker Model Runner:
docker model run hf.co/tcmofashi/mai_replier-9b-v2p1
mai_replier / replier-9b-v2p1 — LoRA Adapter
MaiBot「回复器(replier)」的 9B LoRA adapter,v2p1 数据代际训练。 本仓库只包含 LoRA adapter(465MB),不包含基座模型。 使用方式:挂载到
Qwen/Qwen3-9B(或同结构基座)上即可。
一、这是什么
这是 MaiBot 生态中的「回复器」模型 —— 一个跑在 QQ 群里、扮演具体角色人设发言的群聊 bot。它读取折叠成单条 user 消息的群聊上下文,输出一句符合人设的自然口语化回复。
| 项 | 值 |
|---|---|
| 基座 | Qwen3.5-9B(bf16,混合架构:24 层 linear_attention + 8 层 full_attention) |
| 微调方式 | LoRA SFT |
| LoRA 配置 | r=64, alpha=32, dropout=0.1, target={q,k,v,o,gate,up,down}_proj |
| 训练数据 | 38,956 条 guided(v2p1 数据代际,单轮契约折叠格式) |
| 训练完成度 | 1 epoch(断点续训到 ckpt-1601 收敛) |
| 训练框架 | transformers 5.17 Trainer + accelerate(海光 DCU gfx936) |
| 最终 loss | 收敛(详见训练日志) |
| 角色 | 当前线上 9B 主用模型(vLLM served-name qwen3.5-9b-v2p1) |
二、🚨 关键:v2p1 接受的输入格式 —— 单轮契约折叠
这个模型只接受「单轮契约折叠」格式。 喂多轮 messages 会得到严重退化输出(这是经过实测得到的血泪教训)。
2.1 单轮契约折叠是什么
把整段群聊上下文(system + 多条群友消息 + planner 指引)折叠成一条 user 消息,模型只学在这条 prompt 后接一句话:
<|im_start|>system
{人设 system:人设描述 + 通用模板}
<|im_end|>
<|im_start|>user
{契约 prompt:头部 + 对话行 + 目标行 + [guidance块] + 尾部指令}
<|im_end|>
<|im_start|>assistant
{target 回复}
<|im_end|>
2.2 完整 prompt 模板(训练时使用的格式)
System 部分(人设 + 通用模板,参数化角色):
你的名字是{角色名},也有人叫你{昵称}。
{人设描述(一段人物侧写:性格、口癖、爱好、身份等)}
你正在一个 QQ 群里,需要以这个角色的身份参与群聊。请严格遵守以下要求:
1. 回复要短,符合群聊风格,通常一句话(不超过50字)
2. 不要使用任何 markdown 格式、列表、标题
3. 不要重复使用标点符号(如"!!"、"……"),不要复读
4. 不要复述上下文里别人已经说过的话
5. 不要以"作为XX"开头
6. 不要使用"@某人",直接说话即可
7. 输出只能是这个角色说的下一句话,不要任何额外解释
User 部分(契约折叠后的完整上下文):
你正在qq群里聊天,下面是群里正在聊的内容
其中标注 {角色名}(你) 的发言是你自己的发言,请注意区分:
当前时间:YYYY-MM-DD HH:MM:SS
{发言人1}:{消息1}
{发言人2}:{消息2}
{角色名}(你):{你之前说过的某句话}
{发言人3}:{消息3}
...
现在{目标发言人}说的:{目标消息}。引起了你的注意。
【回复信息参考】
{planner guidance:对当前对话的分析建议,可能为空}
{尾部指令:根据以上对话和指引,以{角色名}的身份说一句话。不要超过50字。现在,你说:}
2.3 完整 Python 调用示例
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
BASE = "Qwen/Qwen3-9B" # 或本地 Qwen3.5-9B 路径
ADAPTER = "tcmofashi/mai_replier-9b-v2p1"
tok = AutoTokenizer.from_pretrained(ADAPTER)
base = AutoModelForCausalLM.from_pretrained(BASE, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, ADAPTER)
system = """你的名字是小千,也有人叫你千绘莉。
你是一个 17 岁的女高中生,性格活泼开朗,喜欢用网络流行语,偶尔会吐槽。
你正在一个 QQ 群里,需要以这个角色的身份参与群聊。请严格遵守以下要求:
1. 回复要短,符合群聊风格,通常一句话(不超过50字)
2. 不要使用任何 markdown 格式、列表、标题
3. 不要重复使用标点符号
4. 不要复述上下文里别人已经说过的话
5. 不要以"作为XX"开头
6. 不要使用"@某人",直接说话即可
7. 输出只能是这个角色说的下一句话,不要任何额外解释"""
user_prompt = """你正在qq群里聊天,下面是群里正在聊的内容
其中标注 小千(你) 的发言是你自己的发言,请注意区分:
当前时间:2026-10-04 17:30:00
阿明:今天吃啥
阿红:想吃火锅
小千(你):火锅+1
阿明:那晚上走起?
现在阿明说的:那晚上走起?。引起了你的注意。
【回复信息参考】
当前讨论约饭,可以表达积极回应。
根据以上对话和指引,以小千的身份说一句话。不要超过50字。现在,你说:"""
messages = [
{"role": "system", "content": system},
{"role": "user", "content": user_prompt},
]
# ⚠️ 必传 enable_thinking=False(与训练时模板行为对齐)
text = tok.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True,
chat_template_kwargs={"enable_thinking": False},
)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(
**inputs,
max_new_tokens=256,
temperature=0.7,
top_p=0.95,
repetition_penalty=1.15,
do_sample=True,
)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
# → 例如输出:走起走起!我已经开始饿了
三、采样参数(线上同配置)
| 参数 | 值 | 说明 |
|---|---|---|
temperature |
0.7 | 不要调太高,群聊腔已足够 |
top_p |
0.95 | (线上未配置,用 vLLM 默认 1.0 也行) |
max_tokens |
256 | 模型只输出一句话,256 足够 |
repetition_penalty |
1.15 | 防复读(关键,群聊场景必开) |
enable_thinking |
False | 必传,否则 think 未闭合 → 输出格式错乱 |
do_sample |
true | 不要 greedy |
四、🚫 常见错误用法(避免踩坑)
❌ 错误 1:喂多轮 messages
# 不要这样!
messages = [
{"role":"system","content":"..."},
{"role":"user","content":"阿明:今天吃啥"},
{"role":"assistant","content":"火锅+1"},
{"role":"user","content":"阿明:那晚上走起?"},
]
后果:模型从未见过这种格式,会输出思维链、复述上下文、复读等严重退化行为。 正确做法:把多轮历史折叠成一条 user 消息(见 2.2 模板)。
❌ 错误 2:忘记 enable_thinking=False
Qwen3.5 chat template 默认 enable_thinking=True,会输出未闭合的 <think> 块。必须显式传 False(训练时也是传 False,模板插入空闭合 think 块)。
tok.apply_chat_template(messages, ..., chat_template_kwargs={"enable_thinking": False})
❌ 错误 3:省略 system 或改写人设位置
人设必须放在 system 里,user 里的对话行标注 {角色名}(你) 必须与 system 中宣称的角色名一致。
❌ 错误 4:关掉 repetition_penalty
群聊场景模型很容易复读上下文短语,1.15 是血泪调出来的值。
五、训练数据(v2p1 代际)
| 数据源 | 条数 | 说明 |
|---|---|---|
| backbone (QCE) | 主体 | 群聊真实上下文 + 角色历史发言 |
| export0905 | ~20,000 | 9/5 导出补充 |
| online_replay | 3,290 | 线上真实重放(SnowLuma) |
| kimi_distill | 1,164 | Kimi 蒸馏高质量回复 |
| 中性化注入 30% | — | 复制带 guidance 样本并清空 guidance |
| 正则项 ~20% | — | 通用/常识/ACG(防过拟合群聊腔) |
最终训练集:sft_v2p1_final.jsonl,110,931 条(本次 hg_9b_v2p1 实际使用了 38,956 条 guided 子集)。
Schema:
{
"sample_id": "qce_xxx",
"source": "backbone_qce | replay_online | kimi_distill | reg_general_sft",
"split": "train | val",
"system": "(人设 + 通用模板)",
"prompt": "(契约折叠后的单条 user 消息)",
"target": "(要学的回复)",
"image_files": ["xxx.jpg"],
"meta": {}
}
多模态(vision)使用方法:
v2p1 支持多模态输入(图像 + 文本)。多模态数据的排布方式为文本在前,图像在后——完整 prompt 文本作为一个 text part 在前,所有图像 parts 附加在最后(与 MaiBot openai_client 序列化逐字同构):
{
"role": "user",
"content": [
{"type": "text", "text": "你正在qq群里聊天,下面是群里正在聊的内容...\n[15:48:55] 豆豆说:[图片]\n[15:49:02] 豆豆说:这价格有人买???\n...\n现在,你说:"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,/9j/4AAQ..."}},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KG..."}}
]
}
多模态数据处理(训练时):
- 图像缩放:
im.thumbnail((1280, 1280))——预缩到 1280 长边(消除 4K 原图的解码/拷贝开销) - 格式:PIL 探测真实格式(
im.format),jpg→jpeg 归一,支持 jpeg/png/webp - 编码:
im.convert("RGB").save(buf, format=...)→ base64 →data:image/{fmt};base64,{b64} - 图像数量:≤3 张(
image_files[:3],避免 token 爆炸) - parts 格式:OpenAI 兼容
{"type": "image_url", "image_url": {"url": "data:image/{fmt};base64,{b64}"}} - 排布:
content = [{"type": "text", "text": prompt}] + img_parts(文本在前,图像在后) - token 估算:每图约 700 token(
n_img * 700),文本 + 图像 token ≤7500(max_len 8192)
多模态调用示例(OpenAI 兼容协议):
import openai, base64, io
from PIL import Image
client = openai.OpenAI(base_url="http://localhost:8200/v1", api_key="x")
# 图像处理:预缩到 1280 长边,jpg→jpeg 归一,base64 编码
def load_image_part(path):
with Image.open(path) as im:
fmt = (im.format or "PNG").lower()
if fmt in ("jpg", "jpeg"):
fmt = "jpeg"
im.thumbnail((1280, 1280))
buf = io.BytesIO()
im.convert("RGB").save(buf, format={"jpeg": "JPEG", "png": "PNG", "webp": "WEBP"}.get(fmt, "PNG"))
b64 = base64.b64encode(buf.getvalue()).decode()
return {"type": "image_url", "image_url": {"url": f"data:image/{fmt};base64,{b64}"}}
# 多模态调用:文本在前,图像在后(≤3 张)
content = [{"type": "text", "text": user_prompt}]
for img_path in image_files[:3]: # ≤3 张
content.append(load_image_part(img_path))
resp = client.chat.completions.create(
model="qwen3.5-9b-v2p1",
messages=[{"role": "system", "content": system}, {"role": "user", "content": content}],
temperature=0.7, max_tokens=256,
extra_body={
"repetition_penalty": 1.15,
"chat_template_kwargs": {"enable_thinking": False},
},
)
print(resp.choices[0].message.content)
六、vLLM 部署
vllm serve /path/to/Qwen3.5-9B \
--enable-lora \
--lora-modules replier-9b-v2p1=tcmofashi/mai_replier-9b-v2p1 \
--served-model-name qwen3.5-9b-v2p1 \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--port 8200
调用方(OpenAI 兼容协议):
import openai
client = openai.OpenAI(base_url="http://localhost:8200/v1", api_key="x")
resp = client.chat.completions.create(
model="qwen3.5-9b-v2p1",
messages=[{"role":"system","content":system},{"role":"user","content":user_prompt}],
temperature=0.7, max_tokens=256,
extra_body={
"repetition_penalty": 1.15,
"chat_template_kwargs": {"enable_thinking": False},
},
)
print(resp.choices[0].message.content)
七、模型档案信息
| 字段 | 值 |
|---|---|
| Adapter 大小 | 465 MB(safetensors) |
| 训练 token 量 | ~5 亿 token(38,956 条 × ~13k avg tokens × 1 epoch) |
| 训练耗时 | ~14h(双卡海光 DCU gfx936) |
| 最终 eval loss | 见训练日志 |
| 训练框架版本 | transformers 5.17, peft 0.20, torch 2.10 |
八、相关资源
- 训练完整手册(数据流、契约渲染、组装脚本):见 MaiBot 主仓
docs/v2p1_single_turn_training.md - MaiBot 主项目:https://github.com/MaiM-with-u/MaiBot
- air_reading 契约折叠插件:负责线上把多轮消息折叠成契约 prompt
九、License
Apache-2.0(与基座 Qwen3.5-9B 一致)
十、免责声明
本模型为角色扮演群聊 bot 用途训练,仅用于模拟具体角色人设发言。请遵守 QQ 平台规则与当地法律法规,不要用于违法、侵权、骚扰等场景。模型输出不代表作者观点。
- Downloads last month
- 26