wear

A small convolutional classifier over Fashion-MNIST clothing images: 28 by 28 grayscale in, one of ten classes out. Two convolution layers (1 to 32 channels, 32 to 64, 3 by 3 kernels) with max pooling, then a 128-unit dense layer with dropout and a 10-way head. 421,642 parameters, trained for five epochs with Adam at learning rate 1e-3, batch size 128, seed 0. Test accuracy 0.9073. CPU training took minutes.

The classes are T-shirt/top, Trouser, Pullover, Dress, Coat, Sandal, Shirt, Sneaker, Bag and Ankle boot, in the dataset's canonical order.

No transformer here. The PreTrainedModel wrapper exists only so the weights serialize as config.json plus model.safetensors and load through AutoModel, the same arrangement as pole. There is no attention, no pretraining, and no transfer story: it classifies small grayscale clothing images and nothing else.

Usage

import torch
from transformers import AutoModel
from hf_wear import WearCNN  # registers the architecture

model = AutoModel.from_pretrained("harpertoken/wear")
model.eval()
image = torch.rand(28, 28)
print(model.predict_label(image))

predict_label takes a 28 by 28 float tensor in range 0 to 1, or a batched 1 by 28 by 28, and returns the class name. Preprocessing is torchvision.transforms.ToTensor() on the raw image and nothing else: no normalization, no augmentation at inference, matching training exactly.

Predictions

First test occurrence of each class, with model predictions

The first test occurrence of each class in evaluation order, with the model's prediction, the true label, and the test index. All ten are correct, which is what the selection rule produced rather than a curated set; overall test accuracy is 0.9073, so roughly one in eleven predictions elsewhere is wrong.

Training

Canonical Fashion-MNIST via torchvision, 60,000 train and 10,000 test, seed 0 throughout. Per-epoch test accuracy ran 0.8696, 0.8855, 0.8888, 0.8978, 0.9073. The wrapper was checked for exact equivalence against the trained module: identical predictions on 2,000 test images in eval mode, after catching that dropout made train-mode outputs differ.

No dataset is published alongside this model. Fashion-MNIST already exists canonically and re-hosting it would add a duplicate, so the card cites the source instead.

Limitations

Small grayscale images of centered clothing items. Anything else, color photos, off-center subjects, classes outside the ten, is out of scope and the scores will be meaningless. 0.9073 is an ordinary result for this architecture, not a benchmark claim.

Downloads last month
31
Safetensors
Model size
422k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results