
BNP Paribas Cardif
Multimodal RAG with Verifiable Visual Evidence Grounding
- Stage
- 4-6 mois
- Hybride
- Rémunéré
- Échéance : 20 oct. 2026
- Artificial Intelligence
- Statistics and machine learning
- natural language processing
- Large Language Models
- Retrieval-Augmented Generation
- Computer vision / 3D scanning
- Computer Vision / Multimodal AI
Description du poste
Context
The NLP team at BNP Paribas Cardif explores advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to improve the reliability, transparency, and efficiency of conversational AI systems used for internal employee assistance and customer support.
Traditional RAG pipelines retrieve text passages before generating an answer. Enterprise documents such as contracts, reports, forms, and financial statements also contain essential information in tables, charts, figures, layout, and typography. Text extraction can lose these visual and structural signals, causing retrieval errors and incomplete or poorly grounded answers.
This internship aims to design and evaluate an end-to-end multimodal document RAG system that produces accurate answers together with verifiable visual evidence. The work will compare text-based and multimodal retrieval approaches, investigate grounded answer generation, and study the trade-off between performance and latency.
Objectives and responsibilities
- Conduct a literature review on multimodal document retrieval, vision-language models, visual evidence attribution, agentic multimodal RAG, and latency-efficient inference.
- Build an end-to-end prototype combining a multimodal document embedder, a multimodal reranker, and a vision-language reader that generates answers with evidence bounding boxes.
- Compare a text-embedding RAG baseline with the multimodal pipeline across plain text, tables, figures, and charts.
- Evaluate retrieval quality, answer quality, and evidence localisation, with particular attention to answers that are both correct and supported by the correct document region.
- Explore adaptation techniques such as LoRA and synthetic data generation to improve retrieval, reranking, and grounded generation across document types, domains, and languages.
- Communicate progress and findings clearly and rigorously to the project team.
- If results are positive, contribute to a reproducible prototype and a potential scientific publication.
- Optionally investigate reinforcement-learning approaches for grounded generation, agentic document navigation, and the latency-performance trade-off of the complete pipeline, including inference time, pages processed, model calls, index size, and GPU memory consumption.
Location and work mode
The position is based in West Paris, Rueil-Malmaison, with a hybrid setup including 50% remote work.