INSPECTION SUBJECT // AI ENGINEER
CURRENT · TEAM LEAD, SMART MFG

Po-Han
Huang

AI Research Engineer · Computer Vision for Smart Manufacturing

I build computer-vision systems that turn factory-floor video into decisions — from sub-5-second defect inspection on live production lines to operator-action recognition for SOP compliance. My work sits exactly where image/video perception meets language models and the physical world.

HOW I MAP TO THE ROLE

From perception to physical systems

Three axes the role cares about — and the concrete work I bring to each.

01

Image & Video Perception

Most of my work lives in pixels and frames. I've shipped multi-view >4K defect inspection at ~5s end-to-end, and built video pipelines that segment work cycles and recognize operator actions from fixed factory cameras.

Defect DetectionTemporal Action Seg.Cycle-Time
02

VLM, LLM & Agents

I've run early feasibility and failure-case studies for VLM-based anomaly detection (e.g. Qwen-VLM) on synthetic PCB scenarios — mapping where today's vision-language models break, and where they're already useful to assist inspection.

VLM / LLMFew-ShotFailure Analysis
03

Toward Physical AI

Manufacturing is where perception meets the physical world. My systems close the loop from camera to production line — verifying SOP compliance, flagging anomalies, and feeding station-level throughput analytics back to operators.

SOP ComplianceHuman-in-the-LoopDeployment
SELECTED WORK

Readouts from the line

Production-grade systems and research, measured by what they changed.

~5s
End-to-end inspection latency
Explainable surface-defect inspection
Deployed an explainable laptop surface-defect pipeline across three manufacturing lines, handling multi-view >4K image batches and streamlining exception handling and verification.
60%
Defect-labeling demand reduced
Diffusion-based synthetic data pipeline
Built a generative pipeline to address the cold-start problem, improving coverage across defect types while meeting production QC constraints.
WACV 26
First-author · training-free
PatchEAD — few-shot anomaly detection
A training-free few-shot framework unifying industrial visual prompting with robust defect localization under limited supervision. Published at WACV 2026.
Video
Cycle-time + motion / SOP
Production-line video understanding
Frame-similarity signals with a hysteresis state machine segment work cycles for throughput analytics; a motion-recognition prototype verifies key operator actions for SOP compliance.
VLM
Feasibility & limits study
Qwen-VLM on synthetic PCB
Early-stage feasibility study and failure-case analysis for VLM-based anomaly detection, outlining key limitations, sensitivities, and improvement directions.
99.9%
Reference-mark detection · <5% error
Two-stage size estimation (intern)
Reached ~92% segmentation mIoU and kept size-estimation error within 5% via data augmentation, ensembling, and stable optimization.
TRAJECTORY

Path so far

May 2024 — Present
AI Research Engineer · Team Lead, Smart Manufacturing
Inventec Corp. — Taipei, Taiwan
Leading computer-vision & AI applications across anomaly detection, cycle-time, and motion/SOP compliance on production lines.
Apr — Oct 2023
AI Research Engineer Intern
Inventec Corp. — Taipei, Taiwan
Two-stage size estimation; co-authored ICASSP 2024 work on data-scarce medical segmentation.
Jan 2022 — Aug 2023
Research Assistant
Academia Sinica — Taipei, Taiwan
Cross-dataset deepfake detection with difficulty-aware training and soft labels. Published at MM Asia 2023.
2017 — 2023
M.S. & B.S., Computer Science & Information Engineering
National Taiwan University of Science and Technology (NTUST)
M.S. thesis on face forgery detection (MM Asia 2023). Consistent Top 1–2% in AI competitions.
RESEARCH & IP

Publications & patents

WACV 2026 · First author
PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection
P.-H. Huang, J.-L. Li, P.-H. Huang, M.-C. Chang, W.-C. Chen
ICASSP 2024 · Equal contribution
Improving Limited Supervised Foot Ulcer Segmentation Using Cross-Domain Augmentation
S.-J. Kuo*, P.-H. Huang*, C.-C. Lin, J.-L. Li, M.-C. Chang
MM Asia 2023 · Equal contribution
Multi-Task Self-Blended Images for Face Forgery Detection
P.-H. Huang*, Y.-H. Han*, E. Chu, J.-C. Chen, K.-L. Hua
Neural Networks 2023 · Journal
DEFAEK: Domain Effective Fast Adaptive Network for Face Anti-Spoofing
J.-D. Lin, Y.-H. Han, P.-H. Huang, J. A. Tan, J.-C. Chen, M. Tanveer, K.-L. Hua
Published Application · US
Segmentation model training method, device, and non-transitory computer readable storage medium — US 2025/0166357 A1, 2025.
In Submission · ×3
Generative defect synthesis · foundation-model-based few-shot learning · explainable defect inspection pipeline.
TOOLKIT

Calibration

CV — Image / VideoDefect & anomaly detection, temporal action segmentation
GenerativeDiffusion models for synthetic data
VLM / LLMFew-shot, prompting, feasibility analysis
LanguagesPython (primary), C / C++
FrameworksPyTorch, OpenCV, Hugging Face
Industrial ADAnomalib (PatchCore / PaDiM), MVTec-style benchmarks
DeploymentONNX, TensorRT export; Netron for graph inspection
ServingFastAPI, Gradio, Docker
WebJavaScript, hand-built static sites, Gradio UIs
InfraLinux, Git, Slurm / HPC
DomainSmart manufacturing, industrial QC