Unsupervised NLP detection of recurring topics
NLPUnsupervisedMLGDPR
ARI 0.88
clustering quality
F1 0.94
on labelled corpus
43 000
real emails processed
Problem
Map the recurring topics in the inbound email flow of a payroll / social-management firm, with no existing label set.
Constraint
Unsupervised learning (no ground truth), ultra-sensitive payroll data (GDPR), and the need to scale.
Approach
- Pipeline designed end to end: ingestion (Microsoft Graph API) → GDPR anonymisation → embeddings → clustering → LLM naming → dataviz.
- Self-hosted BGE-M3 embeddings, UMAP reduction + HDBSCAN clustering, cluster naming by LLM.
- Anonymisation with Presidio + spaCy FR + checksum detectors (NIR, IBAN, SIRET, card numbers).
Result
- Model quality validated on a synthetic labelled corpus: ARI 0.88 / F1 0.94.
- Scaled to ~43,000 real emails, with a macro → micro topic hierarchy and reproducible runs (disk caches).
Stack
BGE-M3UMAPHDBSCANspaCyPresidioMicrosoft Graph
Next project
Valloire Habitat · Multi-source territorial real-estate scoring