Building AI infrastructure that operates at enterprise scale.
6+ years shipping production ML systems — 6B+ API calls/year, 99.9% uptime, 3M+ requests/day.
# Unified Inference Gateway# 3M+ requests/day · EKS · Multi-modelclass InferenceGateway:def __init__(self):self.models = ["llm-30b", "llm-8b", "asr-v3",]self.autoscale = Trueself.sla = 0.999def serve(self, request):# Route to optimal modelreturn self.route(request)
Production systems processing billions of requests with enterprise-grade reliability.
Enterprise AI systems built for scale, reliability, and measurable business impact.
Privacy infrastructure enabling AI to safely process sensitive client data at scale
1.5B API calls · 99.9% availability
Built and scaled an enterprise PII redaction service processing 1.5B API calls in Q1 2026 with 99.9% availability and zero reported breach events — enabling multiple AI products to safely operate on sensitive client interaction data.
Call intelligence platform processing tens of millions of calls for actionable business insights
30.1M calls · 316M minutes processed
Co-designed and scaled an enterprise speech analytics platform processing 30.1M calls and 316.29M minutes of audio, with 30% cost reduction and 25-40x inference speedup through platform modernization.
Unified API layer standardizing access to all enterprise AI capabilities
20+ teams · 100% API delivery
Architected a unified AI SDK providing standardized access to LLM, PII redaction, and embedding services — adopted by 20+ product teams with 100% API delivery, dramatically reducing integration time.
Self-healing responsible AI system that makes models both safer and smarter
Self-healing AI governance
Architected a modular Responsible AI framework combining real-time guardrails with automated self-healing — detecting policy violations per-request and closing the model quality loop via automated LoRA/PeFT fine-tuning.
Production LLM serving framework enabling 10+ enterprise AI use cases
10+ use cases · 3M+ requests/day
Contributed to enterprise LLM enablement including custom model configurations for vLLM serving with PagedAttention, achieving ~2x throughput improvement and enabling cost-effective serving of proprietary models at enterprise scale.
Full-stack ML engineering from model development through production serving and governance.
Client App → API Gateway → PII Service (EKS) → NER Model (PyTorch) → Redacted Output
Audio Ingestion → Denoising → ASR → PII Redaction → Diarization → Classification → Insights API
Product Teams → Client SDK → Auth (Okta) → Service Router → [LLM | Redaction | Embeddings]
Request → RAI Orchestrator (FastAPI) → NeMo Rails Engine → [Block | Pass | Flag] → Evaluation Pipeline → Self-Heal (LoRA)
Request → Inference Gateway (EKS) → vLLM (PagedAttention) → Model Router → [30B | 8B | SLM] → Response
Progressive ownership from ML pipelines to enterprise AI governance.
Aug 2019 - Sep 2021
The Vanguard Group
Foundation — ML infrastructure, inference serving, data pipelines
Sep 2021 - Sep 2024
The Vanguard Group
Scale — production systems processing millions of daily interactions
Sep 2024 - Dec 2025
The Vanguard Group
Depth — multi-model inference, cost optimization, platform unification
Dec 2025 - Present
The Vanguard Group
Governance — production-scale Responsible AI with closed-loop improvement
IEEE Paper
Bansal, A., Joshi, R., et al. · IEEE AAIML 2026 — February 2026
Read on IEEE Xplore →Vanguard IT Division (2025) — Quarterly award selected by a panel of past recipients from all IT division nominations.
NYU Tandon (2018) — MS Computer Science · GPA 3.77/4.0
(2018) — Built "Feed a Homeless" — real-time platform connecting food donors with nearby homeless individuals.
Hackathon Mentor · Vanguard
LLM Integration Patterns · Dec 2023
1-2 engineers annually · mentee achieved promotion