Text-to-Image
Diffusers
lora
stable-diffusion-xl
stable-diffusion-3.5
flow-matching
style
matchbox
india

Indian Matchbox Art Style: Text-to-Image Generator

Two style LoRAs that generate new vintage Indian matchbox labels from a text prompt. Both were trained on the same 197 labels (spearb0lt/Indian-Matchbox-Labels) with the trigger word phlmx:

folder base model objective file
sdxl/ Stable Diffusion XL 1.0 noise prediction (DDPM) phlmx_style_sdxl_ep10_diffusers.safetensors (diffusers) and phlmx_style_sdxl_ep10_kohya.safetensors (ComfyUI, AUTOMATIC1111)
sd3.5-medium/ Stable Diffusion 3.5-medium flow matching pytorch_lora_weights.safetensors (diffusers)

Each folder also has CARD.json: every setting and measured number of its training run. The training notebooks, with all their outputs, are on GitHub: spearb0lt/Indian-Matchbox-Art-Style-Text-to-Image-Generator.

Samples

SD3.5-medium LoRA

SD3.5-medium LoRA samples

SDXL LoRA

SDXL LoRA samples

The peacock (MAYUR) and the swan in this gallery come from an earlier SDXL LoRA run of notebook 55 that was trained on Cooper Hewitt museum prints, whose weights are not published; the other eight images come from the published SDXL LoRA. Each image's run is recorded in images/image_prompts.json.

Every prompt, seed and setting is in images/image_prompts.json.

Which one to use

Left: SDXL LoRA. Right: SD3.5-medium LoRA. Same subject and headline, each with its own notebook's settings. The SDXL peacock comes from the earlier SDXL run of notebook 55 (trained on museum prints, weights not published, seed 478163327 at strength 0.8). More pairs are in images/compare/.

  • SD3.5-medium LoRA: headlines are usually spelled correctly, with flat, saturated colour and clean outlines. Best for legible, bold labels.
  • SDXL LoRA: a muted palette with ink-on-paper texture and engraved line work, like a worn original; headlines are often misspelled. Works in the widest range of tools (kohya file).

Part of the difference in look comes from the base models: with the LoRA switched off (scale 0.0), base SD3.5-medium already draws flat, saturated colour and base SDXL a softer, painted picture; on a short headline such as MAYUR both base models are roughly right. SD3.5 also has a T5-XXL text encoder and joint text-image attention, and its LoRA was trained on hand-written captions that spell out each label's headline. The GitHub README has the full comparison.

Model details

Base models

SDXL 1.0 SD3.5-medium
repository stabilityai/stable-diffusion-xl-base-1.0, with the VAE madebyollin/sdxl-vae-fp16-fix stabilityai/stable-diffusion-3.5-medium (gated: accept the licence first)
denoiser U-Net, 2 567 463 684 parameters MMDiT transformer, 24 blocks, 2 469 663 936 parameters
text encoders CLIP ViT-L (123 060 480) + OpenCLIP ViT-bigG (694 659 840) CLIP ViT-L (123 650 304) + OpenCLIP ViT-bigG (694 659 840) + T5-XXL encoder (4 762 310 656)
how text conditions the image cross-attention joint attention: text and image tokens in one attention per block
VAE 83 653 863 parameters, 4 latent channels 83 819 683 parameters, 16 latent channels
whole pipeline about 3.47 B parameters about 8.13 B parameters
training objective noise (ε) prediction flow matching (velocity prediction)
licence CreativeML Open RAIL++-M Stability AI Community License

Parameter counts are read from each repository's safetensors files.

The LoRA adapters in this repository

sdxl/ sd3.5-medium/
files phlmx_style_sdxl_ep10_diffusers.safetensors (46.62 MB, 1 120 tensors, fp16) and phlmx_style_sdxl_ep10_kohya.safetensors (46.70 MB, 1 120 fp16 weights + 560 fp32 alpha values) pytorch_lora_weights.safetensors (47.84 MB, 486 tensors, fp32)
trainable parameters 23 224 320 (about 0.9 % of the U-Net) 11 943 936 (about 0.5 % of the transformer)
rank / alpha 16 / 16 (LoRA scale 1.0) 16 / 16 (LoRA scale 1.0)
adapted layers to_q, to_k, to_v, to_out.0 in every U-Net attention block: 560 layers image stream to_q, to_k, to_v, to_out.0 and text stream add_q_proj, add_k_proj, add_v_proj, to_add_out: 243 layers
initialisation Gaussian lora_A, zero lora_B, no dropout same
trigger word phlmx phlmx
strength 1.0 (recommended by the run's strength sweep) 0.9 used for the samples; 1.0 scored highest in the run's strength sweep
embedded metadata trigger word, base model, rank, alpha and recommended strength in both files; the diffusers file also holds the PEFT LoraConfig, so diffusers restores rank and alpha exactly the PEFT LoraConfig
training precision bf16 base model, fp32 LoRA parameters, saved as fp16 bf16 base model, fp32 LoRA parameters, saved as fp32
loads in diffusers (load_lora_weights); the kohya file is the format AUTOMATIC1111 and ComfyUI read diffusers (load_lora_weights)
run card sdxl/CARD.json sd3.5-medium/CARD.json

Usage

import torch
from diffusers import StableDiffusionXLPipeline, StableDiffusion3Pipeline, AutoencoderKL

REPO = "spearb0lt/Indian-Matchbox-Art-Style-Text-to-Image-Generator"
NEG = "photograph, photorealistic, 3d render, blurry, low quality, watermark"

# SDXL LoRA
vae = AutoencoderKL.from_pretrained("madebyollin/sdxl-vae-fp16-fix", torch_dtype=torch.float16)
sdxl = StableDiffusionXLPipeline.from_pretrained("stabilityai/stable-diffusion-xl-base-1.0", vae=vae,
                                                 torch_dtype=torch.float16, variant="fp16").to("cuda")
sdxl.load_lora_weights(REPO, subfolder="sdxl", weight_name="phlmx_style_sdxl_ep10_diffusers.safetensors")
img = sdxl("phlmx, a matchbox, a white swan swimming in front of a red sunburst, yellow background, "
           "headline 'SWAN BRAND'", negative_prompt=NEG, width=1152, height=896,
           num_inference_steps=30, guidance_scale=6.0).images[0]

# SD3.5-medium LoRA (accept the base model's licence on the Hub first)
sd3 = StableDiffusion3Pipeline.from_pretrained("stabilityai/stable-diffusion-3.5-medium",
                                               torch_dtype=torch.bfloat16).to("cuda")
sd3.load_lora_weights(REPO, subfolder="sd3.5-medium", weight_name="pytorch_lora_weights.safetensors")
img = sd3("phlmx, a matchbox label, a roaring tiger's head inside a yellow circle, red background, "
          "headline 'TIGER'", negative_prompt=NEG, width=768, height=1152, num_inference_steps=28,
          guidance_scale=5.0, joint_attention_kwargs={"scale": 0.9}).images[0]

Prompts. Start with the trigger and follow the shape of the training captions: phlmx, a matchbox label, <what is shown>, <background colour>, headline '<one or two words>'. Short English headlines render best. Style words such as "vintage" are not needed.

Settings used for the samples. SDXL: LoRA scale 1.0, 30 steps, CFG 6.0, sizes of about 1 megapixel (832 x 1216, 1216 x 832, 1024 x 1024 and similar). SD3.5-medium: LoRA scale 0.9, 28 steps, CFG 5.0, 768 x 1152 or 1152 x 768.

ComfyUI / AUTOMATIC1111: use sdxl/phlmx_style_sdxl_ep10_kohya.safetensors with SDXL 1.0 at weight 1.0.

Training

SDXL LoRA SD3.5-medium LoRA
data 189 train, 8 held out 189 train, 8 held out
captions Florence-2 auto-captions, style words removed (won a four-way caption ablation) the dataset's hand-written captions
resolution 1024-class aspect-ratio buckets 768 x 768 pixel area
LoRA rank 16, alpha 16, attention to_q/k/v/out.0, 560 layers, 23.2 M parameters rank 16, alpha 16, joint attention (image and text streams), 243 layers, 11.9 M parameters
objective ε-prediction, min-SNR-5 weighting flow matching (target ε − x₀), logit-normal timesteps, shift 3.0
schedule 5 670 steps, batch 1, AdamW lr 1e-4, cosine, 30 views per image 1 418 steps of 4 images, lr 1e-4, 30 views per image
selection best of 10 epochs by held-out CLIP style score, prompt adherence and a DINO memorisation check: epoch 10 final step
hardware Colab NVIDIA L4, bf16, 124 min, peak 12.6 GiB Colab NVIDIA L4, bf16, 74 min, peak 13.8 GiB

Measured on held-out evaluation prompts (each column on its own prompts, so not comparable across columns):

SDXL LoRA SD3.5-medium LoRA
CLIP style score, base to LoRA 0.583 to 0.672 0.580 to 0.592
CLIP prompt adherence, base to LoRA 0.321 to 0.320 0.326 to 0.327
generations that copy a training image 0 not measured

Limitations

  • Lettering. Small secondary text is usually invented pseudo-text; the SDXL LoRA also misspells many headlines. Indian scripts (Devanagari, Tamil and others) are not rendered as real words by either model.
  • The trigger can leak into the image. The word phlmx sometimes appears printed on the label.
  • The SDXL captions. Many of its auto-captions were longer than SDXL's 77-token limit, and some kept <pad> tokens, which weakened what it learned about lettering (documented in the notebook, §5.2).
  • Narrow domain. Trained only on matchbox labels; other subjects come out as matchbox labels.

Licence

  • The SDXL LoRA is a derivative of SDXL 1.0 and follows the CreativeML Open RAIL++-M licence.
  • The SD3.5-medium LoRA is a derivative of SD3.5-medium and falls under the Stability AI Community License.
  • The training images are historical commercial prints of mixed provenance whose rights are not cleared (see the dataset card). The LoRAs are shared for research and education.

Credits

The project started from a Reddit post about a "desi-max" style LoRA trained on Qwen-Image (yenupam/desi-max). Its data was not published; these LoRAs use their own data and smaller base models.

Downloads last month
-
Inference Examples
Examples
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for spearb0lt/Indian-Matchbox-Art-Style-Text-to-Image-Generator

Adapter
(113)
this model

Dataset used to train spearb0lt/Indian-Matchbox-Art-Style-Text-to-Image-Generator