LAKEWRIGHT self-hosted

Self-hosted · Your cloud · Your models

Data ingestion that never sees your data.

Lakewright is an AI agent that inventories, classifies, and lands your messiest sources into governed Apache Iceberg tables — running entirely inside your environment. No control plane. No phone-home. Your credentials never leave the building.

Apache-2.0 core Air-gap tested No BAA required
Your environmentmessy filesai answer
NO CONTROL PLANE
Nothing to phone home to. Lakewright brings no cloud of its own.
RUNS IN YOUR VPC
Your object store, your Iceberg catalog, your LLM endpoint.
AUDITED BY BUILD
Every AI decision recorded with its model and prompt, from run one.
HUMAN AT THE GATE
The agent proposes; a person approves; only then does it load.
01 / THE OLD WAY

You already know how this story goes.

Files land on an SFTP share. An expensive integration engine fumbles them into user-facing folders. Somewhere between the drop and the dashboard, trust quietly dies.

Files land, nobody's watching

Vendors drop extracts on an SFTP share on their own schedule, in whatever layout this month's export produced. Renamed columns arrive unannounced.

The integration engine earns its license fee

A six-figure ETL suite processes the drop through jobs nobody remembers writing — then fumbles the output into user-facing shares as more files.

No master data, no memory

The same customer exists four ways across five extracts. Nothing reconciles identity, nothing versions schemas, nothing remembers what was loaded.

Validation is a person squinting at Excel

Quality control is someone opening the dump, eyeballing row counts, and forwarding it on. Analysts hand-build dashboards on mindless Excel exports, trying to derive meaning from data nobody governs — and they don't scale past the person who made them.

One bad file, and the house of cards falls

A date format shifts. A header disappears. The overnight batch dies at 2am, every downstream report is silently wrong by morning, and engineers spend the evening — and the next one — reconstructing what happened from job logs. The executives stop trusting the numbers. That's the real cost.

Incident — night of the 14thevery shop, eventually
22:47vendor_extract_v2_FINAL(3).csv lands on sftp share
23:10nightly batch #4413 picks it up
23:12WARN  column count 41 != 43 — continuing anyway
23:58FATAL job DW_LOAD_MASTER step 17 of 62
00:04retry 1 of 3 … retry 2 of 3 … retry 3 of 3
00:31downstream: 9 jobs skipped, reports built on stale data
07:55CFO dashboard shows Tuesday's numbers. Nobody knows yet.
09:12"quick question — do these revenue figures look right to you?"
21:40engineer, evening two, still diffing job logs by hand

None of this is a people problem. It's what happens when ingestion has no memory, no gate, and no audit — and every fix is another manual step on the pile.

02 / SECURITY

Nothing crosses the line.

Lakewright replaces the pile — the watcher, the validators, the run-books — with one governed engine. And unlike the SaaS tools that promise the same, it does it without your data ever leaving your walls.

Runs where your data lives Self-hosted

Deploy the signed container in your VPC or on-prem. Lakewright reads your stores, calls the LLM endpoint you configure, and writes to your Iceberg catalog. There is no hosted tier for your bytes to transit — because there is no hosted tier.

No BAA, by design HIPAA

Detection, redaction, and loading of PHI all happen in-process, inside your walls. There is no third-party processor touching protected health information — so there is no business associate agreement to negotiate. The compliance surface you simply don't add.

Audited to the decision SOC 2 ready

Every classification, schema proposal, and PII flag is written to a durable record with the exact model and prompt that produced it. The audit trail is a structural property of the engine, not a feature bolted on for the security review.

A human holds the switch Governed

The agent inventories and proposes; a person reviews a declarative plan and approves it; only an approved plan can execute. Autonomy where it's safe, a hand on the switch where it counts.

03 / WORKFLOW

Point it at a store. It takes it from there.

Give it read-only credentials and Lakewright walks the whole path — from an unsorted bucket to governed tables your warehouse can query — pausing for the one decision a human should own.

01ScanA read-only inventory of any store: feeds identified, schemas profiled, PII and PHI flagged, drift and duplicates surfaced. No writes anywhere.
02PlanThe agent proposes declarative ingestion plans — target schema, output shape, PII actions — that you read and edit as plain YAML.
03ApproveThe human gate. A person reviews and approves the plan. Nothing loads that a human hasn't signed off on — enforced in the engine, not just the UI.
04LoadGoverned Apache Iceberg tables. PII actions applied, hashing keyed, every load an audited snapshot linked to the plan that produced it.
05WatchAn ever-running process keeps feeds fresh as new files land, and tails Postgres and MySQL change data as rows move — into an append-only changelog.
04 / COVERAGE

Every signal in. One governed copy out.

Not just clean CSVs. The awkward, regulated, decades-old formats real enterprise data lives in — detected by evidence, classified, and loaded into Iceberg your whole stack already knows how to read.

Tabular & fixed-width
CSV / TSVFixed-widthNACHA / ACHBAI2Excel · ODS
Semi-structured
JSONNDJSONXMLFHIRGeoJSON
Healthcare EDI & clinical
X12 837 / 835 / 834HL7 v2DICOM · PHI
Columnar → Iceberg
ParquetAvroORC*
Documents & OCR
PDF + OCRImagesDOCXEmail + attachments
Change data capture
Postgres logicalMySQL binlog→ changelog
One governed copy · every consumer

Every source lands once in Apache Iceberg. Then your whole stack reads that one governed copy — dbt models it, Snowflake and Databricks query it, Trino and Athena federate it, and your feature store serves it to ML. Open format, open catalog, no lock-in — and archives and codecs (gzip, zstd, tar, zip) expand transparently on the way in.

05 / THE NEW WORLD

Same question. Forty seconds. With receipts.

Remember the 9:12am message from the old world? Here's what it looks like when the data underneath can actually be trusted — and an AI assistant can reach it.

Questions replace dashboards

The dashboards that never scaled? You mostly stop building them. People ask questions in plain language, and their AI assistant answers from the same governed tables — no ticket, no backlog, no three-week wait.

The assistant finds the right data on its own

Every table landed with its meaning attached — what it is, what its columns mean, what's sensitive. So an assistant connected over MCP (the open protocol AI tools use to reach data) finds "the returns data" without anyone teaching it your table names.

It can only see what it's allowed to see

Names, SSNs, card numbers — masked or hashed before the data ever landed. There is no raw copy sitting behind the answer for an AI to leak, because you approved what landed and nothing else exists.

Every answer shows its work

Where the data came from, when it landed, and that it verified — attached to the answer itself. When the number looks surprising, you check the receipts in seconds. That's the thing the old world could never give you: trust that survives a hard question.

The same question — new world9:12 am
boss"quick question — do these revenue figures look right?"
agentchecking the revenue data… (3 files landed overnight, all verified)
agentthe numbers are right — the dip is real.
Northeast returns doubled on the 14th.
boss"which vendor?"
agentAcme Corp — their return file landed 6:02am, 2,140 rows, verified against source.
customer names in that file? masked before landing — I never saw them.
elapsed: 40 seconds · no ticket · nobody's evening

This is what AI in the enterprise actually needs — not a smarter model, but data it can find, trust, and be trusted with. Lakewright builds that ground truth.

06 / POSTURE

Built for the audited enterprise.

Designed from the first commit to fit the programs regulated data teams already live inside — because the compliance conversation is where these tools usually die.

HIPAASOC 2GDPRPCI-DSS
The structural advantage

Because your data and your models never leave your environment, the vendor-risk review is of a self-hosted product, not a data processor. No data-processing agreement to sign, no BAA to negotiate, no fourth party to add to your SOC 2 scope. Governed erasure honors deletion requests against immutable raw capture. Compliance-aware by construction — not a certification we resell, an architecture you can audit.

The wright & the octopus

A wright is a maker — shipwright, wheelwright. The Lakewright builds and tends your lakehouse: arms in every store, a brain in every arm, adapting to whatever it grabs.

Bring it to your data.