Where PhDs and companies meet
Menu
Login

INT-26036-IN RETRIEVAL-AUGMENTED GENERATION FOR FINANCIAL REGULATION VIA LLM ORCHESTRATION AND HYBRID SYMBOLIC-SUBSYMBOLIC AI

ABG-140067 Master internship 6 months See Job Description
2026-08-22
Logo de
Luxembourg Institute of Science and Technology
Belval Luxembourg
  • Computer science

Employer organisation

The Luxembourg Institute of Science and Technology (LIST) is a leading Research and Technology Organisation (RTO) that drives innovation for the economy and society in Luxembourg and beyond. With cutting-edge expertise in Natural, Built, Industrial environments, Space, AI, Security and defence technologies. LIST bridges scientific excellence and applied research to design solutions that address real-world challenges and create positive impact.

Do you want to know more about LIST? Check our website: https://www.list.lu/

 

Your LIST benefits

  • An organization with a passion for impact and strong RDI partnerships in Luxembourg and Europe that works on responsible and independent research projects
  • Sustainable by design, empowering our belief that we play an essential role in paving the way to a green society
  • Innovative infrastructures and exceptional labs occupying more than 5,000 square metres, including innovations in all that we do
  • An environment encouraging curiosity, innovation and entrepreneurship in all areas
  • Personalized learning programme to foster our staff’s soft and technical skills
  • Multicultural and international work environment with more than 60 nationalities represented in our workforce
  • Diverse and inclusive work environment empowering our people to fulfil their personal and professional ambitions
  • Gender-friendly environment with multiple actions to attract, develop and retain women in science
  • 32 days’ paid annual leave, 11 public holidays, 13-month salary, statutory health insurance

Description

Internship contract | Belval | 6 Months

Are you passionate about research? So are we! Come and join us

How will you contribute?

You'll be working on a production-grade Retrieval-Augmented Generation (RAG) system that processes regulatory documents and serves as a critical tool for compliance analysis. You will contribute to create a multi-layered system with real performance requirements and measurable impact

Code Development (60%):

  • Enhance the retrieval algorithm by implementing adaptive weighting based on query characteristics
  • Add comprehensive document validation pipeline with format detection, quality scoring, and automatic preprocessing
  • Implement advanced chunk optimization strategies.
  • Build evaluation frameworks to systematically measure retrieval performance
  • Optimize vector database indexing pipeline for faster document ingestion

Research & experimentation (25%):

  • Test different embedding models against the current Nomic Embed Text baseline
  • Explore cross-lingual capabilities for multilingual EU regulatory documents
  • Research and prototype graph-based retrieval for document relationships
  • Experiment with multi-agent systems and workflow orchestration platforms.

System integration (15%):

  • Enhance the Streamlit interface with advanced filtering and search capabilities
  • Integrate new LLM backends and optimize prompt engineering
  • Implement comprehensive logging and monitoring for production deployment

Profile

Is Your profile described below? Are you our future colleague? Apply now!

Education

  • Advanced degree (Master’s or Bachelor’s final year) in Computer Science, Data Science, Machine Learning, NLP, Information Retrieval, or related quantitative field with strong ML coursework

Experience and skills

Must Have:

  • Strong Python programming (you'll be reading and modifying 1000+ lines of existing code)
  • Understanding of vector spaces, cosine similarity, and basic ML concepts
  • Experience with at least one ML framework (scikit-learn, transformers, etc.)
  • Familiarity with data structures and algorithms
  • Git proficiency for collaborative development

Highly valued:

  • Previous work with LangChain, or similar retrieval systems
  • Experience with NLP libraries (spaCy, NLTK, transformers)
  • Understanding of information retrieval concepts (BM25, TF-IDF, relevance scoring)
  • Experience with Streamlit or similar web frameworks
  • Knowledge of evaluation methodologies for ML systems

Language skills

  • Fluency in English (and French), both oral and written. Other relevant languages are an asset.

Starting date

Dès que possible
Partager via
Apply
Close

Vous avez déjà un compte ?

Nouvel utilisateur ?