Company: Institut Mines-Télécom
Country: France
Type: Onsite
Employment: Entry
Description: Grande general engineering school of IMT-Institut Mines-Télécom, the leading group of engineering schools in France, IMT Atlantique aims to support transitions, train responsible engineers and put scientific and technical excellence at the service of teaching, research and innovation. The position is based in Brest, within the BRAIN (BRoader Artificial INtelligence) team at Lab-STICC (UMR CNRS 6285). The team works on the foundations of machine learning (representations, frugal learning, foundation models, signal processing) and their applications, particularly in health. She recently developed REVE, a foundation model for electroencephalography pre-trained on over 25,000 subjects (NeurIPS 2025), and is interested in discrete diffusion language models and mechanistic interpretability. The ENDIVE project studies what diversity can bring to machine learning when the budget is on the number of annotated examples rather than computing power. In this regime, what a new example brings no longer depends only on its own quality, but on what distinguishes it from those already seen. The technical entry point of the project is sampling with guarantees of diversity, and in particular the determining point processes (DPP), whose kernel matrix encodes both the relevance of the points and their similarity. The project explores this question on two complementary levels: the diversity of the data, i.e. the choice of examples that we annotate, preserve or present to the model; and the diversity of representations, i.e. the choice of descriptors, contexts and models that we exploit or combine. The first results concern the diversified decoding of discrete diffusion language models and the localization of information in the representations of transformers. The proposed position concerns the second level, approached by the tools of mechanistic interpretability. Project page: https://bastienpasdeloup.github.io/endive/ MISSIONS The main missions of the position are as follows: Define and evaluate diversity criteria in the space of representations of foundation models, relying on the tools of mechanistic interpretability. 2. Study the link between the diversity of training data and the diversity of internal mechanisms that the models acquire. 3. Contribute to the scientific production and dissemination of the project. ACTIVITIES: Diversity criteria in representations: • Train and analyze parsimonious autoencoders on model activations of foundation, in order to extract interpretable descriptors. • Construct similarity kernels between these descriptors and evaluate, by means of decisive punctual processes, whether a diverse subset provides more compact coverage than individual importance alone. • Measure the effect of these criteria on downstream tasks in a poorly annotated regime, and compare the network depths at which the representations are read. 2. Diversity of data and diversity of mechanisms: • Compare the descriptors found by parsimonious autoencoders trained on models fed from different data regimes, and quantify their recovery. • Evaluate the stability of these descriptors from one training to another, in order to distinguish the reproducible mechanisms from optimization artifacts. • Relate these measurements to the curation procedures studied in the project, in particular for discrete diffusion language models. 3. Scientific production and project life: • Write and submit the results obtained in international conferences and journals. • Publish the code and experimental protocols necessary for the reproducibility of the results. Level of training and/or minimum experience required: 🎓 Doctorate obtained less than 3 years
Apply here:
Web: Apply here
Emails:
Found 6 similar Onsite jobs