Cloudastructure Inc
Remote
Director of Machine Learning
Aug 2022 – Present
- Own end-to-end delivery of all ML features on the platform, from design and development to production maintenance.
- Built the three production GPU video pipelines (object tagging, face recognition, vehicle analytics) on Kubernetes, extending inference from one GPU architecture to custom builds for Ampere (8.6), Ada (8.9) and Blackwell (12.0).
- Built and led a high-performing ML team through ~100x growth in data volume, optimizing the ML pipeline to scale video processing from 100K to 12M daily.
- Scaled the vehicle pipeline to ~1,630 videos/min on one Blackwell GPU through TensorRT, zero-copy GPU decode, a GIL-free C++ decode feeder with GOP skipping, cross-video overlap and batching.
- Led cloud-to-colocation migration, reducing infrastructure costs by 75% while maintaining uptime and data integrity.
- Own the on-prem GPU cluster (40+ GPUs): built DCGM, Prometheus and Grafana monitoring for utilization, health and per-GPU Xid errors, and eliminated Xid 31 GPU memory faults caused by NVDEC surface exhaustion and a CuPy/PyTorch allocator race.
- Integrated a human-in-the-loop AI-assisted annotation workflow, improving model accuracy through iterative feedback.
- Leading AI acceleration efforts: building coding-agent harness skills and workflows that raise code development standards, and task-specific LLM harnesses for analysis that cut token cost and improve results with smaller models.
Machine Learning Engineer
Nov 2020 – Aug 2022
- Set up the active learning pipeline end to end: data collection, sample selection for labeling and model retraining.
- Designed and implemented a robust model version control pipeline, ensuring reproducibility and seamless tracking of model changes.
- Automated ML operations (MLOps), enabling scalable retraining and efficient model orchestration.

















