PhD: 3D-Aware Multi-Object Tracking
At IDEMIA, we aim to offer our employees, a dynamic and exciting environment where you have opportunities for career growth and professional development.
Internal mobility is a great way for you to energize your career and to build your personal brand. It’s also a great way to explore other functions, professions or countries as IDEMIA operates in different businesses and in 50+ locations around the world.
Purpose
IDEMIA and Télécom Paris are opening a CIFRE PhD position to address a key challenge in modern multi-object tracking (MOT): how can we design more robust tracking systems in the presence of severe and prolonged occlusions? While current tracking methods have achieved impressive performance, occlusion remains one of the main sources of identity switches, missed detections, and fragmented trajectories, particularly in crowded and unconstrained environments. We are seeking to develop an improved tracking system by exploring the hypothesis that 3D-aware and amodal representations can provide the missing information required to maintain object identity through occlusion. The PhD will investigate how reasoning about the full, potentially occluded extent of an object (amodal tracking) and/or its position and structure in 3D space (3D-aware tracking) can be integrated into a tracking architecture.
Particular attention will be given to challenging real-world scenarios involving dense crowds, frequent occlusions, and complex interactions. While the initial focus will be on single-camera tracking, the development of multi-camera tracking systems is of interest, with the objective of exploiting complementary viewpoints and cross-camera geometric consistency to improve identity preservation. The resulting technologies could support IDEMIA's broader biometric and video-security applications, including public-space surveillance, critical infrastructure protection, smart cities, border control, and access management.
Key Missions
- Investigates and develops robust, efficient, and real-time multi-object tracking (MOT) algorithms, with a focus on end-to-end MOT architectures such as MOTIP2 and pedestrian tracking in challenging occlusion scenarios.
- Explores amodal and 3D-aware representations to enhance tracking performance while addressing the computational and resource constraints of edge devices.
- Designs experimental methodologies and benchmarks to evaluate proposed approaches against state-of-the-art MOT methods and real-world datasets.
- Contributes to the design, implementation, and validation of prototypes, with an initial focus on monocular tracking and future extensions toward multi-camera tracking.
- Keeps up to date with advances in computer vision, multi-object tracking, amodal perception, and 3D scene understanding, identifying and transferring relevant innovations to industrial applications.
- Communicates research findings to the IDEMIA and Télécom Paris research and engineering communities, contributing to scientific publications and papers submitted to leading computer vision and machine learning venues.
Profile & Other Information
-
Solid foundations in deep learning, particularly transformer architectures (attention mechanisms, DETR-style query-based models).
-
Experience with computer vision tasks: object detection, tracking, and/or 3D/geometric vision (depth estimation, neural scene representations).
-
Proficiency in Python and PyTorch.
-
Familiarity with vision-language / foundation models (CLIP-like contrastive models, open-vocabulary detectors) is a plus.
Seize all the opportunities of our fast-paced environment. Expect the unexpected.