Image: recorded sampling stages. Action chunks & particle flow:
illustrative.
AT THE INTERSECTION OFGenerative AI+Computer vision+Robot learning↗
Experience
Where I’ve worked
Generative media at Fynd.
Computer vision at Wobot.
SEP 2023 — PRESENT
Fynd / ShopSense
ML Research Engineer
Image generation, super-resolution & video segmentation
· illustration
Building production generative image and video systems, from
distributed model training to optimized GPU inference.
Multi-GPU fine-tuning of SDXL and Flux.Dev for controllable
generation, using representation alignment, distillation,
and LoRA for efficient adaptation and deployment.
Flow-matching super-resolution, including distillation to a
single-step generator.
Video segmentation, inpainting, and restoration pipelines up
to 4K.
~20Kdaily active users across AI media services
Up to 56%lower latency in foreground removal
More engineering detail +
Curated multimodal training data with VLM-assisted
captioning and filtering. Optimized tiled, 4×
super-resolution inference on NVIDIA L4; TensorRT conversion
reduced latency by approximately 40–45% in that pipeline.
FEB 2022 — SEP 2023
Wobot Intelligence
Computer Vision Engineer I
Retail & drive-through vision
· illustration
Built multi-camera person and vehicle tracking for
drive-through journey analytics and production video
inference.
Detection and tracking pipelines using YOLO, DeepSORT, and
ByteTrack.
Reusable deployment workflows with Docker and NVIDIA Triton.
Reworked multithreaded video processing, reducing CPU
utilization from 90% to 40%.
1,000+cameras supported
250+deployment locations
Selected work
Projects and experiments
Open research, practical experiments,
and the code behind them.
07 PROJECTS
01 / ROBOT LEARNINGFeatured
mini-pi0.
From demonstrations to robot actions.
A compact stack for image-conditioned, flow-matching action
policies. From simulation experiments to a visuomotor policy
deployed on my SO-101 arm.
Assembled and calibrated the arm, collected dual-camera
demonstrations with gamepad IK teleoperation, then trained
and deployed a ViT policy. Hardware demos are qualitative.
End to end
A complete experimentation loop
Demonstration collection, training, rollout evaluation, and
failure diagnostics. Interchangeable action backbones and
vision encoders.
NOISE → IMAGESTEP 7 / 7
Actual samples from the repository.
02 / GENERATIVE MODELING
Flow-based Models
Exploring how noise becomes an image. Conditional flow matching
with optimal transport paths and flexible ODE sampling.
Even when conditional flow-matching paths are straight, the
learned marginal velocity field can curve. The method penalizes
pathwise acceleration through a self-guided finite-difference
estimate, without auxiliary encoders or second-order
autodifferentiation.
Accepted at the Workshop on Structured Probabilistic Inference
& Generative Modeling at ICML 2026.
A little about me
How I like to work
My independent work means getting hands-on with the whole
loop—assembling a robot, collecting demonstrations, training
policies, and understanding where they fail.