Lakewright solutions / brief
Extract scanned PDFs and images into Apache Iceberg with self-hosted OCR
Scanned invoices, faxes, and forms become searchable, structured rows in Iceberg — OCR runs on your own GPU with your own vision model, so documents with PHI or PII never leave your environment.
How does the pipeline decide what needs OCR?
PDF profiling reads page count and text-layer presence. Born-digital PDFs extract text directly; scanned ones render page images for the vision model. Images (TIFF/JPEG/PNG) route straight to OCR. Document metadata (kind, pages, dimensions) lands as inventory tables either way.
The steps
export LAKEWRIGHT_LLM_BASE_URL=http://your-gpu:11434/v1 # your model lakewright scan s3://document-dump # flags needs_ocr per file lakewright approve plans/claim_scans.plan.yaml lakewright load plans/claim_scans.plan.yaml --rest $CATALOG
FAQ
Which models?Any OpenAI-compatible vision endpoint — reference setup is qwen3-vl on a single local GPU box.
Email attachments?.eml/mbox files expand automatically: message inventory plus every attachment routed through the same pipeline.