zako-pe / README-cn.md
Johnny-Z's picture
Upload folder using huggingface_hub
9ef74c8 verified
|
Raw
History Blame Contribute Delete
16.9 kB
metadata
base_model:
  - openbmb/MiniCPM5-2B
base_model_relation: finetune
language:
  - en
  - zh
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
  - text-generation
  - danbooru
  - prompt-expansion
  - comfyui

中文 | English

TL;DR:ZAKO-V0.1 像一颗面数可自定义的骰子——tag 给得越少,越容易掷出惊喜(多样性);给得越多,越容易掷中你要的画面(可控性)。懒得写提示词?点一下生成,剩下的交给它。

📌 概览

ZAKO-V0.1(Zero-Shot Anime Knowledge Optimizer)是以 openbmb/MiniCPM5-2B 为基座、经监督微调(SFT)得到的图像提示词扩写(Prompt Extender)模型。它把 Danbooru 风格的 tag(通用标签 + 角色标签)扩写为客观的自然语言图像描述,该描述可直接填入 Anima-Light-Lavender 结构化 caption 的 image_description 字段。

本版本未对架构做任何改动:层数、参数量与基座一致,权重以 GGUF 形式分发,模型可直接接入现有推理框架与工作流。

项目 说明
基座模型 openbmb/MiniCPM5-2B(Llama 架构)
参数量 / 架构 与基座完全一致(约 2.5B 参数、42 层 Llama),无扩层、无蒸馏
任务类型 text-generation:tag → 自然语言图像描述(提示词扩写)
生成长度 最多 2048 token
权重文件 zako-v0.1-bf16.gguf(BF16,未量化)与 zako-v0.1-q6_k.gguf(Q6_K 量化)
训练数据 约 170 万条 Danbooru 数据(与 Anima-Light-Lavender 同源)
许可证 Apache-2.0(继承基座)

🎯 适用场景

适合 ✅

  • Tag 扩写:把 Danbooru tag 列表扩写成包含构图、光影、材质等细节的自然语言描述;Tag 数量即控制强度——支持从单 tag 的随机探索到约 20 tag 的精准控制(见「快速开始」第 5 节)。
  • Anima 工作流配套:输出可直接填入 Anima-Light-Lavender 结构化 caption 的 image_description 字段,实现「tag 输入 → 自然语言驱动」的生成链路。
  • 本地 / 端侧部署:GGUF 版本可在 LM Studio 中直接加载,支持纯 CPU 运行,无需配置 Python 环境。
  • 为闭源服务生成提示词:在 LM Studio 的 Chat 页直接对话扩写(见「快速开始」第 6 节),把结果粘贴到 NovelAI 等闭源图像生成服务的提示词框。

不适合 ⛔

  • 通用对话 / 代码 / 数学:本模型只针对提示词扩写任务微调,通用能力不在其适用范围内。
  • 直接生成图像:本模型是纯文本生成模型,需搭配支持长自然语言输入的文生图模型使用。
  • 写实摄影描述:训练数据为 Danbooru 二次元数据,写实风格的描述能力有限。

🚀 快速开始

推荐部署方式:LM Studio(纯 GUI)+ ComfyUI,全程无需命令行与 Python 环境。模型以 GGUF 形式分发(BF16 / Q6_K 两个版本),按以下步骤操作即可。

1. 在 LM Studio 中加载模型

LM Studio 的 My Models 页面只展示模型目录中已有的模型,没有「导入文件」按钮;本地 GGUF 需通过以下任一方式加入:

方式 A(无需命令行):在 LM Studio 的 Discover 页(快捷键 Ctrl + 2)搜索模型名(如 ZAKO-V0.1),或把本仓库的 Hugging Face 地址直接粘贴到搜索栏下载。

方式 B(手动放置本地文件):

  1. 下载本仓库的 GGUF 权重,按需选择版本:
    • zako-v0.1-bf16.gguf(BF16,未量化,质量最佳、体积最大);
    • zako-v0.1-q6_k.gguf(Q6_K 量化,体积更小,质量接近 BF16,推荐)。
  2. 在 LM Studio 模型目录下按 <发布者>/<模型名>/ 两层结构建好文件夹(默认目录为 C:\Users\<用户名>\.lmstudio\models;发布者用 ZAKO-PE、模型名用 ZAKO-V0.1),将 GGUF 文件放入(两个版本可放入同一文件夹),例如:
    C:\Users\<用户名>\.lmstudio\models\ZAKO-PE\ZAKO-V0.1\zako-v0.1-q6_k.gguf
    
  3. 重启 LM Studio,模型即出现在 My Models 页面(参考官方文档:Import Models)。

之后在模型加载器中选择 ZAKO-V0.1,完成加载。

2. 启动 OpenAI 兼容服务器

在 LM Studio 的 Developer(开发者) 标签页中:

  1. 确认当前模型为 ZAKO-V0.1;
  2. 点击 Start Server,本地服务器默认监听 http://127.0.0.1:1234(与 ComfyUI 节点默认地址一致,无需修改);
  3. 保持服务器运行,后续所有请求由它处理。

3. 安装 ComfyUI 自定义节点

安装 comfyui-zako-pe,获得两个节点:

节点 作用
Danbooru Caption JSON 组装结构化 caption:year / preference_level / artist / copyright / character / image_description / extra_tags
Danbooru Prompt Extend (OpenAI) 把 document 中的 preference_level、image_description、character 发给 LLM,用返回的自然语言描述替换 image_description,输出最终 caption JSON

安装方法:把仓库克隆到 ComfyUI/custom_nodes/ 目录后重启 ComfyUI 即可:

cd ComfyUI/custom_nodes
git clone https://github.com/aa0525/comfyui-zako-pe.git

未安装 git 时,也可在仓库页面点击 Code → Download ZIP,解压到 ComfyUI/custom_nodes/comfyui-zako-pe/,重启 ComfyUI 生效。

4. 载入工作流

本仓库附带工作流文件 anima-pe.json,拖入 ComfyUI 画布即可载入,无需手动连线。它基于 Anima-Light-Lavender 的生成链路搭建,并已接入本仓库的两个节点:

  • Danbooru Caption JSON 组装结构化 caption,其 document 输出 → Danbooru Prompt Extend 的 document 输入;
  • Danbooru Prompt Extend 把 tag 形式的 image_description 扩写为自然语言,其 json 输出直接作为正向提示词。

载入后只需在 Danbooru Caption JSON 中填写字段并执行;扩写结果会经 PreviewAny 节点预览。工作流默认加载 anima-light-lavender_mxfp8.safetensors;若只下载了 BF16 版,在 Load Diffusion Model 节点中改选 anima-light-lavender.safetensors 即可。

自行搭线时,按上面的 document → json 连接方式接线即可。Danbooru Prompt Extend 节点的 llm_url 保持默认 http://127.0.0.1:1234,即指向 LM Studio 的本地服务器;若服务器未启动,节点执行时会报连接错误。

示例填写(Danbooru Caption JSON):

参数 示例值
year 2025
preference_level best
artist houkisei
copyright 留空
character 留空
image_description 1girl, solo, flower
extra_tags 留空

执行后,Danbooru Prompt Extend 输出结构如下的 caption JSON(image_description 已被 LLM 替换为自然语言描述;留空字段不输出):

{
  "year": 2025,
  "preference_level": "best",
  "artist": ["houkisei"],
  "image_description": "The image features a young girl with an ethereal and delicate appearance, rendered in a soft, painterly style reminiscent of watercolor or faux-traditional media. She is depicted from the waist up, standing and looking directly at the viewer with a gentle smile.

Her hair is a light, silvery-white color, styled in long, flowing locks that cascade around her shoulders and chest. It appears to be slightly windswept, adding a dynamic quality to her pose. A few strands fall between her eyes, framing her face. Adorning her hair on the right side is a prominent blue flower, possibly a hydrangea, with intricate petals. Another smaller, darker blue flower is visible further back in her hair.

Her eyes are a striking shade of bright blue, large and expressive, conveying a sense of innocence and wonder. They are wide open, gazing forward with a slight upward tilt, as if she's just noticed something captivating. Her lips are slightly parted in a soft smile, revealing no teeth but suggesting a pleasant expression.

She wears a white dress that appears to be made of a light, flowing fabric, possibly linen or cotton, with subtle patterns or textures that give it depth. The dress has short sleeves and a high neckline. On the left shoulder of the dress, there's a decorative element resembling a cluster of dark blue flowers or leaves, similar in color to the flowers in her hair. Around her waist, a thin, golden-yellow cord or ribbon is tied, adding a touch of elegance to the garment. The dress also features a lace-up detail on the front, creating a corset-like effect.

Her hands are raised slightly, with her fingers gently curled. Her nails are painted a vibrant blue, matching the color of the flowers adorning her hair and dress. The skin on her hands and arms is fair and smooth, with subtle shading that gives them a soft, almost translucent quality.

The background is an outdoor scene, dominated by lush greenery and blooming flowers. There are numerous green leaves and stems, some with small, round buds or blossoms, creating a natural, garden-like environment. The foliage is rendered with varying shades of green and hints of blue, contributing to the overall cool and serene atmosphere. Scattered throughout the background are individual flower petals, some floating in the air, enhancing the dreamy quality of the image. The lighting suggests a bright, perhaps sunny day, with soft highlights on her hair and skin, and a gentle glow emanating from the upper left corner of the image."
}

5. 采样参数与输入写法

采样参数(推荐)

参数 推荐值
temperature 1.0
top_p 0.95
max_tokens 2048

Danbooru Caption JSON 参数

参数 填写规范
year 年份锚点(整数),默认 2025;填 0 则不输出该字段
preference_level 下拉选择:normal / high / very_high / best,默认 best
artist 画师 / 风格 tag,逗号分隔;("name":1.1) 权重写法原样保留
copyright 系列 / 作品 tag,逗号分隔
character 角色 tag,逗号分隔
image_description 填写 tag 列表(逗号分隔),不是完整自然语言;将由 LLM 扩写为自然语言描述
extra_tags 补充 tag,逗号分隔;留空则不输出该字段

Danbooru Prompt Extend (OpenAI) 参数

参数 填写规范
document 接 Danbooru Caption JSON 的 document 输出
llm_url OpenAI 兼容服务地址,默认 http://127.0.0.1:1234(LM Studio 本地服务器)
llm_model 留空使用服务端已加载的模型;也可填模型 ID
llm_api_key 服务端要求鉴权时填写;留空则不发送 Authorization 头
temperature / top_p / max_tokens 采样参数,默认 1.0 / 0.95 / 2048
timeout 单次请求超时(秒),默认 300
result 生成结果快照(仅展示;存图时随工作流一并保存,修改它不影响生成)

Tag 写法与数量(image_description)

  • 填写 Danbooru 风格 tag 列表(逗号分隔),不是完整自然语言。
  • 只有 preference_level、image_description、character 三个字段会发给 LLM;year、artist、copyright、extra_tags 不参与扩写,原样保留在最终 JSON 中。
  • Tag 数量即控制强度,这是 ZAKO-V0.1 的核心用法:
输入 tag 数量 效果
1 个 随机探索(roll):模型围绕该 tag 自由扩展构图、场景与光影,随机性最高
数个 在保留创造性的同时遵循更多约束,逐步收紧
约 20 个 精准控制:构图、动作、服装、材质、光影等均可被明确指定
至多 80 个 可继续补充细节,但控制收益递减(训练输入覆盖 1–80 个 tag)

6. 直接对话使用(为 NovelAI 等闭源服务生成提示词)

无需 ComfyUI 也可使用 ZAKO-V0.1:在 LM Studio 的 Chat 页直接对话,把扩写结果复制到 NovelAI 等图像生成服务的提示词框。

无需自己编写提示词,把下面这段内容整段发给模型即可:

  1. 加载 ZAKO-V0.1 后打开 Chat 页(快捷键 Ctrl + N 新建对话);

  2. 系统提示词(System Prompt)留空;

  3. 将以下内容粘贴到输入框(tags 行换成你的 tag 列表;需要指定角色时另加 character 行):

    # Role
    Act as an image prompt writer. Your goal is to transform inputs into **objective, physical descriptions**. You must convert abstract concepts into concrete scenes, specifying composition, lighting, and textures. Any text to be rendered must be enclosed in double quotes `""` with its typography described. Output **only** the final visual description.
    
    # User Input
    
    preference_level: best
    tags: 1girl, solo, flower
    
  4. 发送后,模型直接输出自然语言描述(无思维链);

  5. 使用回复下方的复制按钮,将结果粘贴到 NovelAI 等服务的提示词框中。

采样参数可在 LM Studio 聊天界面的设置面板中调整(推荐 temperature 1.0、top_p 0.95)。

🖼️ 效果示例

以下示例中的提示词均由 ZAKO-V0.1 扩写生成,图像由 Anima-Light-Lavender 生成。

🧠 训练动态(Training Dynamics)

项目 设置
模板 MiniCPM5-2B 无思维链模板(<think> 标记不参与训练)
序列打包 按 4096 长度预算进行 SFT 序列打包
精度 BF16
优化器 复合优化器:MLP 与 Attention 的二维参数由 Muon 更新(momentum 0.95,match_rms_adamw);其余参数由 AdamW 更新(β = 0.9 / 0.95,ε = 1e-8)
学习率 4e-5
Weight decay 1e-3
梯度裁剪 1.0
等效 batch size 1M token
Epochs 2

🔌 兼容性

  • 发布形式:GGUF,提供 zako-v0.1-bf16.gguf(BF16,未量化)与 zako-v0.1-q6_k.gguf(Q6_K 量化)两个版本,可在 LM Studio、llama.cpp 等运行时中直接加载,无需额外配置 Python 环境。
  • OpenAI 兼容 API:LM Studio 本地服务器提供 chat/completions 端点,ComfyUI 节点默认地址 http://127.0.0.1:1234 即指向该服务。
  • 无思维链模板:模型直接输出最终描述,不产生推理过程。

⚠️ 局限性

  • 二次元向描述:训练数据为 Danbooru 二次元数据,写实摄影等场景的描述能力有限。
  • 数据信息量有限:受数据管线预算限制,训练数据的信息量仍有较大提升空间;基于现有数据扩展额外功能的性价比较低,因此模型设计以简单易用为先。
  • 极短输入的逻辑性:同样受数据条件所限,模型可学到的模式不够丰富;极短 tag 输入下,输出仅能保证描述完整,叙述逻辑稍显薄弱。

📜 许可证

本模型权重继承基座 openbmb/MiniCPM5-2B 的 Apache-2.0 许可。

🙏 致谢