INT-26036-IN RETRIEVAL-AUGMENTED GENERATION FOR FINANCIAL REGULATION VIA LLM ORCHESTRATION AND HYBRID SYMBOLIC-SUBSYMBOLIC AI
| ABG-140067 | Master internship | 6 months | See Job Description |
| 2026-08-22 |
- Computer science
Employer organisation
Website :
The Luxembourg Institute of Science and Technology (LIST) is a leading Research and Technology Organisation (RTO) that drives innovation for the economy and society in Luxembourg and beyond. With cutting-edge expertise in Natural, Built, Industrial environments, Space, AI, Security and defence technologies. LIST bridges scientific excellence and applied research to design solutions that address real-world challenges and create positive impact.
Do you want to know more about LIST? Check our website: https://www.list.lu/
Your LIST benefits
- An organization with a passion for impact and strong RDI partnerships in Luxembourg and Europe that works on responsible and independent research projects
- Sustainable by design, empowering our belief that we play an essential role in paving the way to a green society
- Innovative infrastructures and exceptional labs occupying more than 5,000 square metres, including innovations in all that we do
- An environment encouraging curiosity, innovation and entrepreneurship in all areas
- Personalized learning programme to foster our staff’s soft and technical skills
- Multicultural and international work environment with more than 60 nationalities represented in our workforce
- Diverse and inclusive work environment empowering our people to fulfil their personal and professional ambitions
- Gender-friendly environment with multiple actions to attract, develop and retain women in science
- 32 days’ paid annual leave, 11 public holidays, 13-month salary, statutory health insurance
Description
Internship contract | Belval | 6 Months
Are you passionate about research? So are we! Come and join us
How will you contribute?
You'll be working on a production-grade Retrieval-Augmented Generation (RAG) system that processes regulatory documents and serves as a critical tool for compliance analysis. You will contribute to create a multi-layered system with real performance requirements and measurable impact
Code Development (60%):
- Enhance the retrieval algorithm by implementing adaptive weighting based on query characteristics
- Add comprehensive document validation pipeline with format detection, quality scoring, and automatic preprocessing
- Implement advanced chunk optimization strategies.
- Build evaluation frameworks to systematically measure retrieval performance
- Optimize vector database indexing pipeline for faster document ingestion
Research & experimentation (25%):
- Test different embedding models against the current Nomic Embed Text baseline
- Explore cross-lingual capabilities for multilingual EU regulatory documents
- Research and prototype graph-based retrieval for document relationships
- Experiment with multi-agent systems and workflow orchestration platforms.
System integration (15%):
- Enhance the Streamlit interface with advanced filtering and search capabilities
- Integrate new LLM backends and optimize prompt engineering
- Implement comprehensive logging and monitoring for production deployment
Profile
Is Your profile described below? Are you our future colleague? Apply now!
Education
- Advanced degree (Master’s or Bachelor’s final year) in Computer Science, Data Science, Machine Learning, NLP, Information Retrieval, or related quantitative field with strong ML coursework
Experience and skills
Must Have:
- Strong Python programming (you'll be reading and modifying 1000+ lines of existing code)
- Understanding of vector spaces, cosine similarity, and basic ML concepts
- Experience with at least one ML framework (scikit-learn, transformers, etc.)
- Familiarity with data structures and algorithms
- Git proficiency for collaborative development
Highly valued:
- Previous work with LangChain, or similar retrieval systems
- Experience with NLP libraries (spaCy, NLTK, transformers)
- Understanding of information retrieval concepts (BM25, TF-IDF, relevance scoring)
- Experience with Streamlit or similar web frameworks
- Knowledge of evaluation methodologies for ML systems
Language skills
- Fluency in English (and French), both oral and written. Other relevant languages are an asset.
Starting date
Vous avez déjà un compte ?
Nouvel utilisateur ?
Get ABG’s monthly newsletters including news, job offers, grants & fellowships and a selection of relevant events…
Discover our members
Nantes Université
ASNR - Autorité de sûreté nucléaire et de radioprotection - Siège
Groupe AFNOR - Association française de normalisation
ONERA - The French Aerospace Lab
Aérocentre, Pôle d'excellence régional
ANRT
Laboratoire National de Métrologie et d'Essais - LNE
Servier
ADEME
TotalEnergies
Généthon
SUEZ
Tecknowmetrix
Nokia Bell Labs France
Institut Sup'biotech de Paris
Medicen Paris Region
Ifremer
