Trustworthy Multimodal Representation Learning for Dynamic State Modeling and Prediction
| ABG-140052 | Sujet de Thèse | |
| 19/08/2026 | Contrat doctoral |
- Informatique
Description du sujet
Context:
The increasing availability of heterogeneous multimodal data offers new opportunities for artificial intelligence to model complex systems and predict the evolution of their states over time. Such data may originate from multiple sources and exhibit different structures, dimensionalities, sampling rates, and statistical properties.
However, conventional machine learning approaches often rely on observations acquired at a single time point or consider a limited number of modalities. In real-world dynamic systems, observations may instead be collected over time, at irregular intervals, and different modalities may not always be simultaneously available, and some observations may be partially or entirely missing [Huang, 2025] [Lin, 2025]. Moreover, each modality may provide only a partial and complementary view of the underlying state of the observed system.
Recent advances in multimodal learning, temporal modeling, and representation learning provide new opportunities to jointly exploit these heterogeneous observations. Rather than processing each modality or observation independently, these approaches aim to learn unified representations that capture complementary information across modalities as well as relevant dependencies over time.
Recent work on irregular multivariate time series has also highlighted the difficulty of jointly modeling temporal and cross-variable dependencies when observations are unaligned and irregularly sampled [Li, 2025].
In this context, a fundamental challenge is to learn a dynamic latent representation capable of characterizing the current state of a system from heterogeneous and potentially incomplete observations, while preserving the information required to model and predict its future evolution.
Problem statement and aim:
The central challenge of multimodal artificial intelligence is no longer merely to combine multiple data sources, but to learn a reliable representation of a dynamic system from observations that are inherently partial, heterogeneous and collected over time.
Recent approaches to incomplete multimodal learning have investigated different strategies, including missing-modality reconstruction, prompt-based adaptation, multimodal generalization and refinement, and deficiency-resistant representation learning [Lang, 2025] [Huang, 2025] [Lin, 2025]. Despite these advances, learning consistent and reliable representations under varying modality availability remains an open challenge. In particular, approaches based on the reconstruction of missing modalities may face difficulties in preserving the reliability and distributional consistency of reconstructed information [Dai, 2026]. These difficulties become even more important when observations are irregular, asynchronous, noisy, or subject to distribution shifts.
Consequently, an important open research question remains: How can a unified dynamic representation be learned directly from heterogeneous and partially observed data while remaining robust, uncertainty-aware, and interpretable?
The main objective of this thesis is to develop trustworthy multimodal representation learning approaches capable of modeling dynamic states and predicting their future evolution from incomplete temporal observations. Rather than treating missing observations solely as a reconstruction problem, the proposed research will investigate representation-centric learning strategies that preserve complementary information across available modalities while adapting to irregular observations and varying modality availability.
Particular attention will be devoted to four complementary properties of trustworthy prediction: robustness to partial observations and distribution shifts, uncertainty quantification, confidence calibration, and temporal explainability.
Approach:
To achieve the stated objectives, the research strategy will be structured around two main phases.
1- Multimodal and Temporal Representation Learning
The first phase will investigate representation-learning approaches for modeling dynamic states from heterogeneous and partially observed multimodal data. The objective is to learn a unified latent representation that captures complementary information across available modalities while preserving relevant temporal dependencies. Rather than systematically reconstructing missing modalities, the proposed research will investigate representation-centric strategies capable of constructing a reliable latent state directly from the observations that are available. Particular attention will be given to preserving shared and modality-specific information, adapting the learned representation to different modality configurations, and limiting representation shifts caused by missing observations. Recent work on incomplete multimodal learning demonstrates the importance of learning robust and consistent representations when modalities are unavailable [Huang, 2025] [Lin, 2025], while more recent approaches further investigate representation consistency under missing-modality conditions [Chen, 2026].
The temporal dimension will be integrated directly into the representation-learning process. In particular, the research will consider irregular sampling and asynchronous observations, for which temporal observations and modalities may not be aligned. Recent work on irregular multivariate time-series modeling confirms that jointly capturing temporal dependencies and dependencies among unaligned observations remains a significant learning challenge [Li, 2025].
The resulting representation should therefore characterize the current latent state from the available observations while retaining relevant information from previous states. This dynamic representation will then support the prediction of future states under varying data-availability and temporal conditions.
2- Robust, Uncertainty-Aware and Explainable Prediction
The second phase will investigate the trustworthiness of the learned representations and their associated predictions under incomplete, noisy, and shifted data conditions. Robustness will be evaluated through controlled perturbation scenarios involving different patterns and levels of missing modalities, input degradation, and distribution shifts.
Beyond predictive performance, uncertainty estimation will be investigated to determine whether the model can identify situations in which the available information is insufficient or unreliable. Recent approaches have highlighted the relevance of explicitly accounting for uncertainty when learning from incomplete multimodal observations [Nguyen, 2025]. Particular attention will be given to predictive uncertainty and confidence calibration, with the objective of ensuring that model confidence appropriately reflects prediction reliability [Cheon, 2026].
The relationship between robustness and uncertainty will constitute an important component of this investigation. A trustworthy model should maintain stable predictions when sufficient information remains available, while appropriately increasing its uncertainty when degradation, missingness, or distribution shift prevents reliable prediction.
Finally, explainability approaches will be investigated to characterize the contribution of individual modalities and temporal observations to the learned representation and resulting predictions. The objective will be to determine not only which information influences a prediction, but also when this information becomes relevant during the evolution of the observed system.
Together, these components aim to establish a general framework for trustworthy dynamic-state modeling and prediction from incomplete multimodal and temporal observations. The specific tasks of this thesis are:
- State-of-the-art analysis and problem formulation: Review and analyze recent approaches in multimodal and temporal representation learning, with particular attention to incomplete modalities, irregular and asynchronous observations, and trustworthy prediction. The main limitations of existing approaches will be identified and the research problems addressed in the thesis will be formalized.
- Dynamic multimodal representation learning: Develop representation-centric approaches for learning unified and dynamic latent representations from heterogeneous and partially observed multimodal data, while preserving complementary information across available modalities and temporal observations.
- Trustworthy modeling and prediction: Develop and evaluate mechanisms for improving robustness to missing modalities, noise, and distribution shifts, while integrating uncertainty quantification, confidence calibration, and temporal explainability into the proposed models.
- Experimental evaluation and validation: Evaluate the proposed approaches on multiple multimodal and temporal datasets under different modality-availability, temporal-irregularity, perturbation, and distribution-shift scenarios, considering predictive performance, robustness, uncertainty calibration, explainability, and generalization.
References:
[Lang, 2025] Lang, J., Cheng, Z., Zhong, T., & Zhou, F. (2025). Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 39(17), 18035–18043. DOI: 10.1609/aaai.v39i17.33984.
[Huang, 2025] Huang, W., Chen, Y., Jiang, X., Gao, C., Zhang, T., Chen, Q., & Wang, Y. (2025). Mitigating Pervasive Modality Absence Through Multimodal Generalization and Refinement. Proceedings of the AAAI Conference on Artificial Intelligence, 39(25), 26796–26804. DOI: 10.1609/aaai.v39i25.34883.
[Lin, 2025] Lin, H., Tang, X., Li, H., et al. (2025). T²DR: A Two-Tier Deficiency-Resistant Framework for Incomplete Multimodal Learning. Findings of the Association for Computational Linguistics: ACL 2025, 8602–8616. DOI: 10.18653/v1/2025.findings-acl.452.
[Li, 2025] Li, B., Luo, Y., Liu, Z., Zheng, J., Lv, J., & Ma, Q. (2025). HyperIMTS: Hypergraph Neural Network for Irregular Multivariate Time Series Forecasting. Proceedings of the 42nd International Conference on Machine Learning (ICML 2025), PMLR 267, 35502–35518.
[Nguyen, 2025] Nguyen, D.A., Do, Q.H., Doan, K.D., & Do, M.N. (2025). Are you SURE? Enhancing Multimodal Pretraining with Missing Modalities through Uncertainty Estimation. arXiv:2504.13465.
[Chen, 2026] Chen, J., Cheng, S., Yutao, Y., Zhang, Y., Yuan, H., Peng, P., & Zhong, Y. (2026). PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities. Proceedings of the AAAI Conference on Artificial Intelligence, 40(24), 20076–20082. DOI: 10.1609/aaai.v40i24.39093.
[Dai, 2026] Dai, R., Jian, A., Zhang, R., et al. (2026). Multimodal Learning with Missing Modalities: Can Progressive Diffusion Achieve Distribution-Consistent Learning from Incomplete Multimodal Data? Knowledge-Based Systems, Article 116272. DOI: 10.1016/j.knosys.2026.116272.
[Cheon, 2026] Cheon, J., & Paik, S.-B. (2026). Brain-Inspired Warm-Up Training with Random Noise for Uncertainty Calibration. Nature Machine Intelligence, 8, 602–613.
Prise de fonction :
Nature du financement
Précisions sur le financement
Présentation établissement et labo d'accueil
Efrei Paris, école d’ingénieurs, composante de l’Université Paris-Panthéon-Assas, est un établissement privé d’enseignement supérieur technique, reconnu par l’Etat, EESPIG, dont dépend le laboratoire Efrei Research Lab, dirigé par Etienne PERNOT.
L’Efrei Research Lab est le laboratoire de recherche de l’Efrei. Il se compose d’une cinquantaine d’enseignants-chercheurs en informatique et électronique ainsi que d’autant de doctorants. Depuis janvier 2022, en intégrant l’université Paris-Panthéon-Assas, Efrei Research Lab est reconnu comme le laboratoire numérique de l’Université, unité de recherche 202224306D, rattaché à l’école doctorale ED 455 EGIC, délivrant le doctorat en informatique.
Ses domaines de recherche se concentrent sur les domaines du numérique à travers quatre axes :
- données et Intelligence Artificielle ;
- sécurité, résilience et confiance numérique ;
- réseaux de communication ;
- systèmes embarqués intelligents.
Le Laboratoire se concentre sur de la recherche appliquée avec deux domaines d’applications majeurs : les sciences du vivant (santé, agriculture et biodiversité, sport, éducation) et les territoires intelligents (entreprises, habitations, réseaux).
L’Efrei Research Lab s’est engagé dans la mise en œuvre de sa responsabilité sociétale vis-à-vis des enjeux environnementaux à travers l’ensemble des activités du laboratoire. Le chercheur de l’Efrei Research Lab s’engage à prendre en compte la transition écologique pour un développement soutenable dans ses activités de recherche menées au sein de l’Efrei Research Lab.
Site web :
Intitulé du doctorat
Pays d'obtention du doctorat
Etablissement délivrant le doctorat
Ecole doctorale
Profil du candidat
- Diplôme : Master’s degree (research-oriented) or an Engineering degree in computer science
- Compétences scientifiques : We are looking for a candidate with a strong background in Machine Learning and Deep Learning.The ideal candidate should have:
- a strong mathematical background, particularly in linear algebra, probability, statistics, and optimization;
- very good programming skills in Python;
- the ability to understand, implement, and critically analyze recent machine learning research papers;
- previous research experience, ideally through a Master’s thesis, research internship, or scientific publication.
- Experience in at least one of the following areas would be particularly appreciated:
- Representation Learning;
- Multimodal Learning;
- Self-Supervised Learning;
- Time-Series or Sequence Modeling;
- Probabilistic Machine Learning
- Langues : Excellent niveau en Français et en anglais.
- Qualités personnelles : bon relationnel pour le travail en équipe, rigueur scientifique, autonomie et esprit d’initiative..
Vous avez déjà un compte ?
Nouvel utilisateur ?
Vous souhaitez recevoir nos infolettres ?
Découvrez nos adhérents
Laboratoire National de Métrologie et d'Essais - LNE
Aérocentre, Pôle d'excellence régional
Ifremer
Généthon
ADEME
ASNR - Autorité de sûreté nucléaire et de radioprotection - Siège
ANRT
TotalEnergies
SUEZ
Servier
Groupe AFNOR - Association française de normalisation
Institut Sup'biotech de Paris
Medicen Paris Region
Nantes Université
ONERA - The French Aerospace Lab
Nokia Bell Labs France
Tecknowmetrix
