
BNP Paribas Cardif
Automatic Prompt Optimization and Reliable Probability Estimation for LLM-based Classification
- Stage
- 4-6 mois
- Hybride
- Rémunéré
- Échéance : 20 oct. 2026
- Artificial Intelligence
- Statistics and machine learning
- natural language processing
- Large Language Models
- Prompt Optimization
- Probability Calibration
Description du poste
Context
BNP Paribas Cardif increasingly relies on Large Language Models (LLMs) for customer request classification and routing, document analysis, intent and complaint detection, summarization, and draft reply generation. Within the Global AI Office, teams design, evaluate, and industrialize AI solutions based on open-weight models deployed on internal infrastructure.
This applied research internship addresses two complementary challenges:
- Automatic Prompt Optimization (APO): Prompt performance can vary significantly depending on wording, examples, and formatting. APO methods automate the search for effective prompts while controlling optimization cost and data requirements.
- Reliable probability estimation: Class probabilities extracted from LLM outputs can be poorly calibrated. Option-token probabilities may become excessively concentrated because of instruction tuning, label or position biases, tokenization effects, and the divergence between option-letter logits and the model’s generated answer.
Objectives
Automatic Prompt Optimization
- Conduct a literature review and select representative methods, including LLM-based search methods such as APE and OPRO, textual-gradient methods such as ProTeGi, evolutionary approaches such as GEPA, and programmatic frameworks such as DSPy/MIPROv2.
- Benchmark the selected methods on classification and generation tasks against the existing evaluation framework.
- Compare the methods in terms of performance gain, optimization cost, sample efficiency, stability, and robustness across models.
Reliable Probability Estimation
- Investigate probability distortion in multiple-choice LLM classification, including label or position bias, tokenization effects, and the divergence between token-level scores and the model’s generated answer.
- Implement and compare probability extraction strategies, including restricted-logit scoring, full-label likelihoods, contextual calibration, permutation-based debiasing, and sampling-based estimates.
- Study the combination of probability extraction and post-hoc calibration techniques such as Platt scaling and temperature scaling.
- Benchmark performance using Expected Calibration Error (ECE), Brier score, log-loss, AUROC for error detection, and selective accuracy/coverage, while monitoring latency and inference cost.
Work environment
As an intern, you will work with public and internal datasets, develop reusable code, analyze and communicate experimental results, and collaborate with the team to integrate the outcomes of the research into the production environment.
Location and work mode
The position is based in West Paris, Rueil-Malmaison, with a hybrid setup including 50% remote work.