I build computer-vision systems that turn factory-floor video into decisions — from sub-5-second defect inspection on live production lines to operator-action recognition for SOP compliance. My work sits exactly where image/video perception meets language models and the physical world.
Three axes the role cares about — and the concrete work I bring to each.
Most of my work lives in pixels and frames. I've shipped multi-view >4K defect inspection at ~5s end-to-end, and built video pipelines that segment work cycles and recognize operator actions from fixed factory cameras.
I've run early feasibility and failure-case studies for VLM-based anomaly detection (e.g. Qwen-VLM) on synthetic PCB scenarios — mapping where today's vision-language models break, and where they're already useful to assist inspection.
Manufacturing is where perception meets the physical world. My systems close the loop from camera to production line — verifying SOP compliance, flagging anomalies, and feeding station-level throughput analytics back to operators.
Production-grade systems and research, measured by what they changed.