Data products / analytics / applied AI

I build data products that turn raw systems into clear decisions.

I design analytics, pipelines, models, and AI products that hold up beyond the demo. The goal is simple: make complex data useful to the people doing the work.

Portrait of Vaibhav Khurana
Vaibhav KhuranaData analysis, engineering, science, and AI
$200Kannual contract removed

Reporting brought in-house at Morgan & Morgan

<15 minreport delivery

Previously three to five days

250K+records per cycle

Governed compliance and compensation pipelines

$100K+policy savings

Hybrid PTO approach adopted by leadership

Selected live work

Built, tested, and available to explore.

A cross-section of deployed work across data engineering, analysis, data science, and AI engineering.

Browse all 12 live projects
01Technical project

Finance & risk

Payments Fraud Risk Data Platform

A governed fraud-risk validation register that publishes all 1,852,394 allowlisted simulated events for bounded analytical queries while keeping identity-like fields, model scores, and payment actions out of the public system.

FocusData Engineering

ResultRetain the ordinary logistic baseline for the measured 1% review queue. It reaches 0.160 PR-AUC and 51.3% recall while the class-weighted challenger performs worse on ranking, recall, and probability error.

PythonPostgreSQLscikit-learnSupabaseFastAPINext.js
02Technical project

Healthcare / Operations

GP Access Planner

A public-data planning product that forecasts recorded general-practice appointments across England while keeping observed access signals and hypothetical capacity explicitly separate.

FocusSenior Healthcare Analyst

ResultDelivered a source-traceable 7, 14, and 28-day planning surface for 104 sub-ICBs, backed by 32.9 million validated source rows, rolling-origin evaluation, immutable releases, and a live edge API.

PythonPostgreSQLdbtscikit-learnLightGBMCatBoost
03Deployed technical prototype

Legal & knowledge

Legal Discovery Graph

Evidence-linked investigation across documents and entities

FocusAI engineer / Data engineer

ResultConnects documents, people, events, and citations in one inspectable investigation workflow.

PythonFlaskLangChainsentence-transformersONNX RuntimePostgreSQL + pgvector
04Technical project

Finance & risk

Automobile-Loan First-EMI Default Strategy Portfolio

A retrospective credit-policy platform over 233,154 Indian vehicle loans: portfolio and vintage reporting, score-decile and segment risk analytics, a policy workbench with editable economics, a per-loan inspector, and post-deployment monitoring.

FocusData Analyst

ResultA credit-policy analyst can test where to draw a first-EMI risk line, see the confidence interval around the answer and who it declines, and take a recommendation or a refusal to governance. On the published assumptions the honest output is a refusal: no evaluated band clears zero.

PythonJavaScriptscikit-learnSQLiteFastAPIReact

How I work

Build the path from question to action.

I define the decision, validate the data, compare a useful baseline, and ship the result in a form someone can actually operate.

  1. Name the decision

    Who acts, what changes, and what a wrong answer costs.

  2. Earn the dataset

    Reconcile definitions, permissions, leakage, and failure paths before tuning anything.

  3. Set the stopping rule

    Choose baselines, capacity, uncertainty, and refusal gates before reading the result.

  4. Ship for review

    Put the recommendation, explanation, and override in the same operating workflow.

Preferred delivery stack

From first brief to monitored decision system.

This is the complete toolkit I would reach for to shape, build, test, deploy, observe, and report on a data product. It is a preferred operating stack—not a usage counter or a dump of project dependencies.

01

Product direction

Briefs, user flows, architecture decisions, and an owned backlog before implementation starts.

Notion
Figma
GitHub
02

Application layer

Accessible interfaces, typed contracts, production APIs, and clear client–server boundaries.

TypeScript
React
Next.js
Python
FastAPI
03

Data foundation

Transactional storage, analytical models, orchestration, event streams, and low-latency access.

PostgreSQL
DuckDB
dbt
Apache Airflow
Apache Kafka
Redis
04

Analysis and modeling

Reproducible exploration, feature engineering, model training, evaluation, and explainable outputs.

pandas
Polars
NumPy
Jupyter
scikit-learn
PyTorch
05

Applied AI

Grounded generation, retrieval, graph context, model access, and portable inference paths.

OpenAI
LangChain
Hugging Face
Neo4j
ONNX Runtime
06

Quality and security

Unit, integration, browser, lint, dependency, and container checks before anything is released.

pytest
Vitest
Playwright
Ruff
ESLint
Snyk
Trivy
07

Cloud and infrastructure

Portable services, declarative infrastructure, managed data, edge delivery, and environment parity.

Docker
Terraform
Microsoft Azure
Supabase
Cloudflare
08

Release and operations

Version control, automated delivery, production hosting, traces, errors, and operating feedback.

Git
GitHub Actions
Vercel
OpenTelemetry
Sentry
09

BI and decision support

Governed metrics, semantic reporting, executive dashboards, ad-hoc analysis, and visual diagnosis.

Power BI
Microsoft Excel
Snowflake
Plotly
Grafana

Career highlights

A short record of outcomes, not a public resume.

As a Data Analyst at Morgan & Morgan, P.A., I build shared data systems, reporting products, and decision models. The profile keeps the public story focused on a few measurable changes.

View achievement snapshots

Start a conversation

Bring me the question that is still difficult to answer.

Discuss a role or project