We are looking for a research intern to join us for a research project aimed at publication at a top-tier venue. The intern will design and develop novel systems that explore the interaction between human pose understanding and vision-language models (VLMs), advancing how these modalities can be combined to reason about human motion, activity, and embodied behavior across images and video.
Our group develops hand and body pose tracking algorithms for various apple devices and applications. One such example includes the hand tracking input for the Vision Pro.