Drop a screenshot of ANY website or app, and THE INSPECTOR — a film-noir detective — works it as a crime scene: he circles each UX flaw on the real pixels, names the charge, and files a verdict with a letter grade. A UX audit that plays like a detective thriller.
But the verdict is just the opening statement. Now it goes further:
⚖️ THE TRIAL — put the interface on trial. The guilty UI elements take the stand and defend themselves while the Inspector rules from the evidence. 🖼️ THE RECONSTRUCTION — one click and FLUX.2 Klein rebuilds the worst element FIXED, live. Before/after, on the real pixels. 🔊 THE VOICE — hear the verdict read aloud (Kokoro, local, no keys). 🚨 MOST WANTED — a public rogues' gallery. Book your case onto a shared board where the city's worst interfaces are ranked by their crimes. Booked by the public.
Three small models, all on Modal (scale-to-zero), none over 32B: 👁️ Qwen2.5-VL-7B (vision agent) · 🖼️ FLUX.2 Klein (reconstruction) · 🔊 Kokoro-82M (voice)
ICYMI, you can fine-tune open LLMs using Claude Code
just tell it: “Fine-tune Qwen3-0.6B on open-r1/codeforces-cots”
and Claude submits a real training job on HF GPUs using TRL.
it handles everything: > dataset validation > GPU selection > training + Trackio monitoring > job submission + cost estimation when it’s done, your model is on the Hub, ready to use