← All projects
Logo IGAM
IGAM ↗

Data project · client

Data

Unsupervised NLP detection of recurring topics

NLPUnsupervisedMLGDPR

ARI 0.88

clustering quality

F1 0.94

on labelled corpus

43 000

real emails processed

Problem

Map the recurring topics in the inbound email flow of a payroll / social-management firm, with no existing label set.

Constraint

Unsupervised learning (no ground truth), ultra-sensitive payroll data (GDPR), and the need to scale.

Approach

  • Pipeline designed end to end: ingestion (Microsoft Graph API) → GDPR anonymisation → embeddings → clustering → LLM naming → dataviz.
  • Self-hosted BGE-M3 embeddings, UMAP reduction + HDBSCAN clustering, cluster naming by LLM.
  • Anonymisation with Presidio + spaCy FR + checksum detectors (NIR, IBAN, SIRET, card numbers).

Result

  • Model quality validated on a synthetic labelled corpus: ARI 0.88 / F1 0.94.
  • Scaled to ~43,000 real emails, with a macro → micro topic hierarchy and reproducible runs (disk caches).

Stack

BGE-M3UMAPHDBSCANspaCyPresidioMicrosoft Graph

Next project

Valloire Habitat · Multi-source territorial real-estate scoring

→