LLM extraction of technical data for industrial quoting
Vision LLMFull-stackEvaluationProduction
85–96 %
measured accuracy
EU
data residency
288
material grades
Problem
Automate industrial quoting by extracting technical data (dimensions, loads, tolerances, materials) from highly heterogeneous documents, to feed the in-house “E8” calculator without manual re-entry.
Constraint
Documents in varied formats, data that must stay within the European Union, materials expressed as free text to be normalised, and a requirement for field-by-field measurable accuracy.
Approach
- Production extraction pipeline with a vision LLM (Gemini, chosen for EU data residency and its ability to read technical drawings).
- Structured, typed outputs (Pydantic), a versioned data schema, and iterative prompt engineering.
- Hardened extraction with fuzzy matching (rapidfuzz) of a free-text material against a catalogue of 288 client grades, plus physical unit conversion (metric ↔ imperial).
- Evaluation system objectively measuring field-by-field accuracy (in-house comparison engine, then Langfuse).
Result
- Extraction accuracy measured at 85 to 96 % depending on the documents.
- Full-stack application delivered: FastAPI / SQLAlchemy async / PostgreSQL, React 19 / TanStack / Tailwind, containerised (Docker, Dokploy), Clerk auth, S3 storage (SHA-256 deduplication), Logfire observability.
Stack
GeminiPostgreSQLFastAPIReact 19LangfuseLogfireDocker
Next project
ODERIS · AI classification of slides at scale for due diligence