Principal ML Engineer · Vanguard

Rajeev Joshi

Building AI infrastructure that operates at enterprise scale.

6+ years shipping production ML systems — 6B+ API calls/year, 99.9% uptime, 3M+ requests/day.

inference_gateway.pyConnected
# Unified Inference Gateway
# 3M+ requests/day · EKS · Multi-model
class InferenceGateway:
def __init__(self):
self.models = [
"llm-30b", "llm-8b", "asr-v3",
]
self.autoscale = True
self.sla = 0.999
def serve(self, request):
# Route to optimal model
return self.route(request)

Impact at Scale

Production systems processing billions of requests with enterprise-grade reliability.

0B+
API Calls Processed
Annual production volume — PII Service
0M+
Call Hours Processed
Call Intelligence Engine — annual volume
0M+
Requests/Day
Unified Inference Gateway
0%
Cost Reduction
Platform Modernization

Signature Projects

Enterprise AI systems built for scale, reliability, and measurable business impact.

InfrastructurePrivacyAI/ML

Enterprise PII Redaction

Privacy infrastructure enabling AI to safely process sensitive client data at scale

1.5B API calls · 99.9% availability

Built and scaled an enterprise PII redaction service processing 1.5B API calls in Q1 2026 with 99.9% availability and zero reported breach events — enabling multiple AI products to safely operate on sensitive client interaction data.

AI/MLInfrastructurePlatform

Call Intelligence Engine — Enterprise Speech Analytics

Call intelligence platform processing tens of millions of calls for actionable business insights

30.1M calls · 316M minutes processed

Co-designed and scaled an enterprise speech analytics platform processing 30.1M calls and 316.29M minutes of audio, with 30% cost reduction and 25-40x inference speedup through platform modernization.

PlatformInfrastructure

Enterprise AI SDK

Unified API layer standardizing access to all enterprise AI capabilities

20+ teams · 100% API delivery

Architected a unified AI SDK providing standardized access to LLM, PII redaction, and embedding services — adopted by 20+ product teams with 100% API delivery, dramatically reducing integration time.

AI/MLPlatform

Responsible AI Framework

Self-healing responsible AI system that makes models both safer and smarter

Self-healing AI governance

Architected a modular Responsible AI framework combining real-time guardrails with automated self-healing — detecting policy violations per-request and closing the model quality loop via automated LoRA/PeFT fine-tuning.

AI/MLPlatform

Enterprise LLM Platform

Production LLM serving framework enabling 10+ enterprise AI use cases

10+ use cases · 3M+ requests/day

Contributed to enterprise LLM enablement including custom model configurations for vLLM serving with PagedAttention, achieving ~2x throughput improvement and enabling cost-effective serving of proprietary models at enterprise scale.

Technical Depth

Full-stack ML engineering from model development through production serving and governance.

Enterprise PII Redaction

Client App → API Gateway → PII Service (EKS) → NER Model (PyTorch) → Redacted Output

Call Intelligence Engine — Enterprise Speech Analytics

Audio Ingestion → Denoising → ASR → PII Redaction → Diarization → Classification → Insights API

Enterprise AI SDK

Product Teams → Client SDK → Auth (Okta) → Service Router → [LLM | Redaction | Embeddings]

Responsible AI Framework

Request → RAI Orchestrator (FastAPI) → NeMo Rails Engine → [Block | Pass | Flag] → Evaluation Pipeline → Self-Heal (LoRA)

Enterprise LLM Platform

Request → Inference Gateway (EKS) → vLLM (PagedAttention) → Model Router → [30B | 8B | SLM] → Response

Career Arc

Progressive ownership from ML pipelines to enterprise AI governance.

Aug 2019 - Sep 2021

Machine Learning Engineer

The Vanguard Group

  • Built multiprocessing inference pipeline — 85M+ rows, 70% execution time reduction
  • Designed event-driven DB architecture for speech platform
  • Created NLP labeling platform and EDA acceleration package
  • Productionized Next Best Action model for advisor application

Foundation — ML infrastructure, inference serving, data pipelines

Sep 2021 - Sep 2024

Senior Machine Learning Engineer

The Vanguard Group

  • Scaled Call Intelligence Engine to 30.1M calls and 316M minutes processed
  • Built Client SDK adopted by 20+ product teams
  • Adapted enterprise LLM experimentation platform (RAG, DynamoDB, Pinecone)
  • Built PII service to production scale

Scale — production systems processing millions of daily interactions

Sep 2024 - Dec 2025

Senior Machine Learning Engineer

The Vanguard Group

  • Co-developed unified inference gateway — 3M+ requests/day
  • Contributed vLLM custom configs with PagedAttention (~2x throughput)
  • Co-designed speech analytics processing 300K+ calls/day
  • Architected LLM SDK with shared prompt library and cross-team reuse

Depth — multi-model inference, cost optimization, platform unification

Dec 2025 - Present

Principal Machine Learning Engineer

The Vanguard Group

  • Architected RAI — Responsible AI with self-healing model loop
  • Drove red-teaming protocols and automated safety scoring
  • Designed AI-ready data pipeline with automated LoRA/PeFT fine-tuning
  • Built framework-agnostic guardrails (NeMo + LangGraph + Mem0)

Governance — production-scale Responsible AI with closed-loop improvement

Publications & Recognition

IEEE Paper

Cost-Efficient, Model-Agnostic, Low-Latency LLM Deployment

Bansal, A., Joshi, R., et al. · IEEE AAIML 2026 — February 2026

Read on IEEE Xplore →

IT Peer Recognition Award

Vanguard IT Division (2025) — Quarterly award selected by a panel of past recipients from all IT division nominations.

Academic Achievement Award

NYU Tandon (2018) — MS Computer Science · GPA 3.77/4.0

MLH HackNYU — Third Place

(2018) — Built "Feed a Homeless" — real-time platform connecting food donors with nearby homeless individuals.

Women in Data Science

Hackathon Mentor · Vanguard

FAST Hackathon

LLM Integration Patterns · Dec 2023

Engineer Mentorship

1-2 engineers annually · mentee achieved promotion

Let's build something significant.

Open to principal/staff engineering roles, and research collaboration.

Philadelphia, PA

click me · drag me