Data & Platform Engineer · AI/ML Infrastructure on Kubernetes · Fail-Closed Data & Agentic Systems — New York

Hossain
Pazooki

I build data platforms, evaluation harnesses, and reliable agentic & LLM systems for regulated domains — point-in-time financial data, healthcare claims, cross-border compliance — where a wrong answer has a real cost and "it usually works" isn't good enough.

What the work proves

Correctness you can audit — not confidence you have to trust.

34% fewer
claim denials, after a production failure exposed systematic LLM over-confidence and I rebuilt the decision gate.
context · Orawell Group — AI/ML Tech Lead, healthcare RCM
five-gate
verification stack — schema, semantic, source-span, conflict, expert-attestation — decides what the model is allowed to act on.
source · ATLAS — verification stack
fail-closed
on unevaluable, not just on failure. Absence of valid proof is a denial, never a default-pass.
concept · Intent Layer — authorization
86.6M
gold facts
point-in-time-correct, distilled from 181M bronze rows across 16,667 filers. An as-of date returns only what was already public.
source · VANTAGE — point-in-time
Selected work

A portfolio of one conviction.

A policy plane that compiles regulation into signed, executable artifacts, applied ML held to its own evidence, and a point-in-time data platform. Data and models interpret; deterministic engines execute; every decision leaves a trace.

policy plane
regulatory engine

ATLAS regulatory-rule-engine ↗

A regulatory rule-execution engine, and the policy plane of a three-repo system. Compiles contradictory multi-jurisdiction regulation into content-addressed, signed, typed-attestation artifacts; a Rust kernel executes them deterministically and fails closed on a verification miss.

A five-gate verification stack — schema, semantic, source-span, conflict, expert-attestation — decides what a rule is allowed to assert. Every artifact is a pure function of its input: a three-language contract test holds the canonical encoding byte-identical across Rust, Python, and WASM, and Rust↔Python execution equivalence is proven over 1,326 generated scenarios. Three downstream consumers re-derive trust independently and refuse a non-published artifact even when its signature is valid.
Rust workspaceBLAKE3 content-addresseded25519-signedtyped attestationsWASM verifier
applied ML
multi-omics QA

CLUE upstream-label-correction ↗

Upstream label correction for multi-omics cohorts. Generates fidelity-checked synthetic cohorts with planted, dial-able label corruption, then catches the mislabeled samples through cross-omics concordance scoring — before bad labels poison everything downstream.

The synthetic cohort is the measuring instrument, not the result: it probes corruption rates real data cannot, and the loop escalates until it finds the hardest rate the detector still clears. That detector now also clears an outside oracle — run unmodified on the real precisionFDA training matrices, at its default threshold, scored against the challenge organizers' own key: F1 0.914. The blind test labels were withheld by the challenge, so that gap is stated rather than closed.
Pythonscikit-learnsynthetic multi-omicsorganizer-keyed oracle
data platform
point-in-time lakehouse

VANTAGE pit-fundamentals-lakehouse ↗

A point-in-time-correct SEC fundamentals lakehouse. An as-of date D returns only what was filed and accepted by D — holding across restatements, so an as-of query never sees a filing that wasn't yet public.

181,351,169 bronze rows distilled to 86,615,392 gold facts across 16,667 filers; 63 of 69 quarters ingested, 6 refused by a fail-closed data-quality gate with recomputed causes. Deployed and published on Databricks serverless under a write-audit-publish rule: audit green on the candidate, the same audit red on a mutated twin, a post-publish consumer probe against a negative control. An Airflow DAG now carries that shape — publish is structurally unreachable when the audit is red or unevaluable — and a dbt layer over gold serves the as-of view with its point-in-time invariants tested in CI, each proven able to fail on a planted twin.
ScalaSparkDelta LakeDatabricks ConnectAirflowdbt-duckdbfail-closed DQ gate
Technical skills

Across the stack, end to end.

Data platforms & pipelines

  • Scala · Spark · Delta Lake
  • Point-in-time correctness & as-of restatement
  • Databricks — serverless + Connect, write-audit-publish
  • ETL, fail-closed data-quality gates, content-addressed lineage

Evaluation & correctness infra

  • Validation harnesses over planted ground truth
  • Fail-closed gates — unevaluable never passes
  • Precision / recall / hallucination tracking
  • Reproducible, signed, content-addressed artifacts

Agentic AI & LLMs

  • Multi-agent orchestration & fan-out builds
  • Hybrid RAG — BM25 + vector retrieval
  • Confidence calibration & verification gates
  • NLI / entailment verification

Languages & systems

  • Python — FastAPI, async, strict typing
  • Rust — deterministic execution kernels
  • Scala / Spark · Go gates · TypeScript / React
  • Kubernetes / EKS · PostgreSQL · CI/CD
Career

Where the conviction was earned.

Nov 2024 — Jan 2026 · Full-time

AI/ML Tech Lead — Orawell Group

New York, NY · healthcare revenue-cycle management
  • Led a team building agentic compliance pipelines — PydanticAI agents orchestrating claim review, denial prediction, and appeal generation with human-in-the-loop checkpoints.
  • Designed a multi-layer confidence-calibration and verification-gate system after a production failure exposed systematic over-confidence in LLM decisions — reduced claim denial rates by 34%.
  • Built a hybrid retrieval pipeline (BM25 + vector) over regulatory and payer-policy documents, with a structured eval framework tracking precision, recall, and hallucination rate.
  • Owned the infrastructure: FastAPI, PostgreSQL, Kubernetes on EKS, Terraform IaC, CI/CD with strict mypy enforcement and 450+ CI-gated tests.
Mar 2022 — Aug 2024 · Full-time

Co-Founder / Lead Engineer — Engager Inc.

New York, NY
  • Built the platform from the first commit to production — establishing architecture, CI/CD, and engineering practice.
  • Designed and shipped the core backend and frontend (React, Python).
  • Owned ingestion and analytics pipelines feeding product and growth decisions.
2020 — 2022 · Full-time

Founding ML Engineer — 3DEO

Los Angeles, CA · metal additive-manufacturing R&D
  • Built the ML platform: a unified data layer linking process parameters, characterization, and production outcomes by part ID across sintering logs, mechanical-test data, and QA.
  • Shipped model-serving hooks that scored new production runs automatically, plus drift monitoring against characterization ground truth.
  • Translated materials-team objectives — yield, faster qualification — into tractable ML KPIs with non-technical domain experts.
2019 — 2020 · Full-time

Data Engineer — Penske Media Corporation

Marketing analytics & revenue
  • Built marketing-analytics models and reporting infrastructure supporting ad sales and audience monetization across PMC's portfolio.
2017 — 2019

Econometrics Research — Institute for New Economic Thinking

Sovereign-wealth-fund investment & capital-allocation models
University of Southern California
B.S. Electrical Engineering · B.A. Economics · 2-yr econometrics research at INET
2013 — 2019