AI classification of slides at scale for due diligence
LLMGDPRArchitectureScale
250 000+
slides (≈ 1,500 missions)
GDPR
embeddings computed locally
cost-aware
LLM arbiter on ambiguous cases
Problem
Capitalise on Vendor Due Diligence reports by automatically classifying slides into a business taxonomy, at the scale of 250,000+ slides (around 1,500 missions).
Constraint
Strict GDPR: the client refuses any data transfer outside the EU. LLM call costs to be kept under control over a massive volume.
Approach
- Cost-aware cascade pipeline: regex → embeddings → LLM arbiter called only on ambiguous cases.
- BGE-M3 embeddings run locally and systematic anonymisation (spaCy + GLiNER) before any call to the vision LLM (Mistral Pixtral, EU).
- Hexagonal architecture; parallelisation via a bounded pool with retry/backoff.
Result
- Automatic classification of slides into the business taxonomy, with the LLM arbiter used only on ambiguous cases to keep cost under control.
- End-to-end GDPR compliance: embeddings computed locally and anonymisation before any LLM call — no client data leaves the EU.
Stack
Mistral PixtralBGE-M3 (local)PostgreSQLpgvectorHexagonal arch.SigNoz
Next project
IGAM · Unsupervised NLP detection of recurring topics