Tech LeapLLC

Services

Four things teams hire me for

Usually several arrive together. The platform needs modernizing, something on it is already broken, and there is an AI initiative depending on both. Each diagram below is interactive: press the buttons and drag the sliders to see the failure and the fix.

AI & ML data enablement

The model is the easy part. The hard part is the supply chain feeding it: training sets that reflect reality, embeddings that refresh when the source changes, retrieval that returns the right context, and monitoring that tells you when it quietly stops working.

I build that layer end to end, covering medallion curation, training-data preparation, embedding generation and vector indexing, inference orchestration, and drift detection backed by an auditable prediction log.

source systems medallion curation training sets vector index inference prediction log ingest curate embed train retrieve log drift detected → rebuild index INDEX ROWS NO LONGER MATCHING SOURCE 0
day 0

Drag the slider. The pipeline is healthy and every job is green the whole way across.

The loop in brass is the part most teams skip. Without it the index silently ages against its source and the model degrades with nothing failing. Rates shown are illustrative.
Databricksvector searchembeddingsRAG retrievalmodel monitoringMLflow

Lakehouse modernization

Medallion architecture, Delta Lake, Unity Catalog migration, and incremental loads that actually stay incremental instead of quietly reverting to full reloads nobody costed.

Also the unglamorous half that decides whether the platform is affordable: cluster right-sizing, job-chain consolidation, and cutting multi-hour runtimes down to something you can schedule a business around.

BEFORE source pipeline target full table truncate + load 4h 10m nightly AFTER source pipeline target changed rows MERGE 35m nightly 6h batch window 7x less compute for the same data
120M rows

Drag it. Full reload scales with the whole table. Incremental scales with the part that changed, which barely moves.

Runtime bars drawn to scale, anchored on a real migration at 120M rows. The saving is not a faster cluster; it is moving only what changed, then keeping it that way under change control.
Delta LakeUnity CatalogPySparkAzure / AWSCI/CDcost tuning

Pipeline rescue

Something has been wrong for months and nobody can say what. Silent zero-record loads. Race conditions in parallel loops. A type promotion that widened a column. A filter hardcoded to a year that has already passed.

I find it, prove it with the data, fix it under your change process, then build the validation that would have caught it in the first week.

CONTROL PLANE scheduler job run status: idle watched DATA PLANE source transform rows written: idle not watched row-count assert > 0 alerting idle signals

Press either button. Same job, same empty source table, two different outcomes for the people who depend on it.

Every green dashboard in the first run is telling the truth. The job did succeed. Nothing in the chain is asking whether it wrote anything.
root-cause analysisdata qualityreconciliationbacklog recoveryaudit readiness

Geospatial data engineering

Spatial data has its own failure modes. Projections that silently disagree, geocoding that drifts, boundary files that stopped matching the addresses joined against them, and volumes that outgrow the single machine they were designed for.

I build enterprise geospatial pipelines covering spatial enrichment, geocoding, feature extraction, reconciliation across point-of-interest, address, boundary and imagery sources, and distributed processing when scale demands it.

DISTRICT A DISTRICT B true location EPSG:3857, unconverted joins to A, correct joins to B, wrong spatial join 12,847 ADDRESSES JOINED TO DISTRICTS 3,412 in the wrong district

Every one of those 3,412 rows loaded clean. Valid address, valid district, no error anywhere in the run.

Neither row errors. Both look like valid addresses with valid districts. The only way to catch it is to reconcile the join against a known-good reference.
PostGISArcGIS / ArcPyGeoPandasQGISspatial indexingGEOINT

How engagements work

Diagnosis before prescription

Most teams can't specify this work up front, because the problem hasn't been diagnosed yet. Asking you to scope it before we look would just move the risk onto you. So we don't start there.

01

Assessment

One to two weeks · fixed fee

I go through the platform, covering pipelines, orchestration, data quality, lineage and cost, and come back with what's broken, what it's costing you, and what to do in what order. You own the document regardless of what happens next. If we never work together again, it still has to be worth what you paid.

02

Build

Fixed scope · fixed price

The assessment names the work and prices it: the migration, the AI data layer, the reliability framework, whatever it turned up. Defined deliverables, defined end date, no open meter. You know what you're buying because we just spent two weeks establishing it.

03

Fractional

Ongoing · 10 to 20 hrs/week

For teams that need senior data engineering continuously but can't justify a full-time hire, or that need someone accountable for the platform between hires. Usually follows a build, because by then I already know the system.

Tell me what's broken.

A short description of the platform and what isn't behaving is enough to start. I'll tell you within a day whether it's something I can help with.