Apply for community grant: Personal project (gpu: A100 Large)

#1
by Ryukijano - opened

@hysts Hi! I'm applying for a community GPU grant for my portfolio Space. Megan from HF support suggested I apply here after I reached out about restoring a previous A10G grant (lost when the Space was accidentally deleted and recreated).

Space: https://huggingface.co/spaces/Ryukijano/CatCon-One-Shot-Controlnet-SD-1-5-b2

I built a single Space that hosts 10 interactive AI demos โ€” surgical phase recognition, quantum error correction, drug repurposing, robotics control, mechanistic interpretability, and more. Think of it as a living lab notebook where each project is a hands-on exhibit, not a static README.

The interesting engineering challenge: how do you run 6+ large models on one GPU without everything catching fire?

The answer is modal inference. Instead of loading every model at startup and watching VRAM evaporate, I built a FastAPI backend that loads models on-demand and frees them when idle. A GPU semaphore controls concurrency. When a user opens the DINO-Endo demo, DINOv2 ViT-14s loads, runs inference, and gets evicted. When they switch to Cosmos Sentinel, Cosmos Reason 2 (8B) takes its place. torch.cuda.empty_cache() is called between swaps so VRAM actually gets returned to the pool โ€” not just "freed" in Python but still hoarded by the allocator.

This isn't just model.to("cuda") and hope for the best. It's a proper lifecycle:

Model registry tracks what's loaded, what's cold, and what's being requested
LRU eviction drops the least-recently-used model when VRAM pressure rises
Lazy imports โ€” heavy dependencies (diffusers, transformers, decord) only import when the relevant demo is activated, keeping cold-start fast
Semaphore-gated GPU access โ€” only one model does forward pass at a time, preventing OOM from concurrent requests
What's in the portfolio:

Demo Model What it does
DINO-Endo Surgery DINOv2 ViT-14s + V-JEPA2 Surgical phase recognition from endoscopic video
Cosmos Sentinel Cosmos Reason 2 (8B) + BADAS Predictive collision detection with natural language risk narration
Syndrome-Net QEC Stim + RL agents Quantum error correction with surface/color codes
Gemma-GR00T Siglip+Gemma+Diffusion head VLA Language-conditioned robotics control
RelP-SAE TransformerLens + SAEs Discovering interpretable features in transformers
ReNova RDKit + ChEMBL Drug repurposing with molecular optimization
Parameter Golf Various compressed models Extreme model compression (4-bit, pruning, distillation)
QuantumForge Rust tensor cores + Python Hybrid quantum-classical computing
Surface Code in STEM Stim + JAX Educational quantum error correction toolkit

Why an A100:
The models range from 4GB (PaliGemma) to 16GB (Cosmos Reason 2 in fp16). With modal inference, peak VRAM usage hits ~20GB when the largest models are active. An A100 40GB gives enough headroom for the model swap pipeline to work smoothly load next model while previous is still flushing, rather than serial load-evict-load that makes demos feel sluggish.

On cpu-basic, none of this runs. The models just sit there like sports cars in a parking lot.

Context:

This Space originally had an A10G community grant from the 2023 Google Cloud x HF Diffusers Sprint (ranked 10th). The Space was accidentally deleted and recreated, which reset the hardware. Megan confirmed I should re-apply for a community grant for the new Space.

Thanks for considering happy to answer any questions!

Best, Gyanateet Dutta (@Ryukijano )

Sign up or log in to comment