TL.

Impact · Applied R&D · Production ML

Work delivered
in real contexts.

4 case studies · Français / English

Four professional case studies connect behavioral research, model evaluation, machine-learning architecture and product decisions.

Descriptions intentionally focus on transferable methods, decisions and learnings. No client data or confidential information is disclosed.
4consolidated case studies
>1Maggregated behavioral observations
Offline → Prodfrom evaluation to system
Evidence + limitsexplicit evidence and limitations

01 — INDUSTRY WORK

Four complete journeys, from hypothesis to product decision.

01
Behavioral AI · Predictive segmentation

Predictive behavioral segmentation: knowing when a model does not generalize

An anonymized case study in behavioral feature engineering, XGBoost, temporal validation and failed transfer across environments.

XGBoostMouse trackingAblationsTemporal validation
Scope
Behavioral modeling
Outcome
Aggregated AUC from 0.70 to 0.77 across behaviors; insufficient cross-domain transfer led to an environment-specific strategy.
02
Multimodal AI · Search & Recommendation

From CLIP benchmarks to a production multimodal platform

A consolidated case study covering model selection, vector retrieval, serving and productionization of a text-image pipeline now running in production.

CLIPClickHouseHNSWFastAPIKubernetes
Scope
Multimodal retrieval
Outcome
Task-specific model selection and an incremental text-image pipeline now in production, with monitoring and load testing.
03
Behavioral AI · Learning to rank

Personalizing product search through behavioral affinities

An anonymized case study covering user-product signals, LambdaMART, temporal validation and the limits of offline evaluation.

LambdaMARTXGBoostNDCGTemporal split
Scope
Ranking personalization
Outcome
Affinities outperform context alone across offline NDCG, Hit Rate and MRR; online uplift remains to be measured.
04
Generative AI · Quantitative UX

Evaluating a vision-language model against human UX judgments

An experiment across 100 ecommerce websites and 600 evaluations comparing GPT-4.1 mini with a public human benchmark on six design-perception dimensions.

GPT-4.1 miniVision-language modelParquetPearsonSpearman
Scope
Multimodal UX evaluation
Outcome
600 evaluations across 100 websites; overall Pearson correlation of 0.416, with positive bias requiring calibration before product use.

Collaborate

A product problem deserves a measurable method.

I work across behavioral signals, search, recommendation and AI systems applied to user experience.

Discuss a need