I build vision‑language models that watch streaming video, hold a compact memory of what's normal, and reason deeply only at the moments that break the pattern. Work spans distributed pretraining, synthetic‑to‑real data pipelines, transformer acceleration, and edge deployment — from first pixel to production verdict.
The project started as a data problem before it was a modeling one. I defined the action taxonomy, recording protocols, camera‑viewpoint specifications, and validation procedures behind a proprietary surveillance‑action corpus, then built the compact model that runs on it: 35,000+ verified clips and 137,000+ pose sequences after geometric augmentation, feeding a model small enough for an edge box.
Multi‑viewpoint capture, cleaning, scene‑level partitioning, controlled augmentation, and a strict held‑out split so no scene ever leaks between train and test.
View‑aware kinematic features feeding a FiLM‑conditioned Temporal Convolutional Network — 824K parameters, sized for Jetson‑class edge hardware.
0.9661 weighted F1 across 27,511 sequences from scenes the model had never seen — evidence that data‑centric work plus a small model can beat a bigger one under an edge budget.
The underlying surveillance‑action dataset stays private — it was built inside a commercial research effort and isn't publicly distributed. A technical artifact package is available on request for research, academic, or hiring evaluation, covering:
Request a copy: ayasin@shrinkhaltai.com
MSc, Artificial Intelligence & Data Science
BSc, Information Systems