Open to full-time AI / ML Engineer roles · US

Raghav Upadhyay

> _

I build production-grade LLM systems — from retrieval to evaluation to deployment.

RAG pipelines with hybrid retrieval and cross-encoder reranking, hallucination detection, LLM-as-a-judge evaluation wired into CI, uncertainty-aware deep learning, and published research on LLM behaviour. M.S. Data Science, University of Arizona (May 2026).

0 AI / ML Projects
0 GPA / 4.00
0 Live LLM App
0 Paper · NeurIPS ’26
SCROLL
// 01 — about

Engineer, not just experimenter

I'm Raghav Upadhyay, an AI Engineer focused on the unglamorous part of LLM work: making systems reliable. I care about retrieval quality, grounded answers, measurable evaluation, and shipping things people can actually use.

My flagship project, AskMyDocs, is a production-grade RAG app with hybrid retrieval, cross-encoder reranking, hallucination detection, and an LLM-as-a-judge harness that fails the CI build when answer quality drops below threshold. Beyond LLMs, I've built uncertainty-aware deep learning for predictive maintenance and co-authored a cross-cultural study of LLM-generated social networks currently under review at NeurIPS 2026.

Lately I've been pushing into vision-driven robotics — building an autonomous maze-navigating agent in Webots that has to find its way using nothing but a first-person camera feed.

Currently finishing my M.S. in Data Science at the University of Arizona (May 2026) and open to full-time AI / ML Engineer roles across the US.

📍 West Lafayette, IN 🎓 MS Data Science · UArizona ⚡ Available May 2026
// 02 — projects

Things I've built & shipped

02
Computer Vision · Robotics ● building now

Vision-Based Virtual Maze Navigator

An autonomous e-puck robot agent that has to find a target inside an unknown 3D maze in the Webots simulator using only its first-person camera — no ground-truth map coordinates. Targeting the full perceive → map → plan → act loop (visual SLAM, path planning, closed-loop control). Live camera-feed processing and keyboard teleoperation are shipped; an HSV target detector is being wired into motor control now.

OpenCV Webots NumPy Robotics
View on GitHub →
03
LLM Research · Publication

LLM-Generated Social Networks — Cross-Cultural Study (Capstone)

Replicated and extended an ICWSM 2025 paper on whether LLMs generate structurally realistic social networks. Built a full generation–analysis–benchmark pipeline across a 4×4×4×3 experimental matrix (prompting × cultures × languages × GPT-4.1 tiers, 96+ verified conditions), quantifying inter-model divergence and identifying political affiliation as the dominant homophily dimension. Written up as a paper now under review at NeurIPS 2026.

GPT-4.1 NetworkX pandas Experiment Design
Read on arXiv →
04
Deep Learning

RUL Prediction with LSTM — Uncertainty-Aware Predictive Maintenance

Benchmarked four LSTM architectures (Vanilla, Stacked BiLSTM, LSTM-Attention, CNN-LSTM) on NASA C-MAPSS for turbofan Remaining Useful Life. Monte Carlo Dropout and Deep Ensembles for uncertainty quantification (mean ± 2σ risk-adjusted decisions), gradient-based XAI with temporal attention heatmaps, and an ipywidgets dashboard with tiered maintenance alerts.

PyTorch TensorFlow Uncertainty XAI
View on GitHub →
05
NLP

Commonsense Reasoning with Pre-trained Language Models

Benchmarked RoBERTa-MNLI and OPT-1.3B on commonsense reasoning datasets, analyzing zero-shot and fine-tuned performance across PIQA and related tasks with Hugging Face Transformers.

RoBERTa OPT-1.3B PyTorch Hugging Face
View on GitHub →
06
ML / NLP

Hate Speech Detection

Text classification pipeline detecting hate speech in online content with Logistic Regression and Random Forest, reaching 92% accuracy on a Kaggle-sourced dataset to support content-moderation use cases.

Python NLP scikit-learn
View on GitHub →
07
NLP / App

Abstractive Text Summarizer

Streamlit app that condenses long articles into short summaries with a distilBART (sshleifer/distilbart-cnn-12-6) seq2seq model, with adjustable length bounds and summarizer logic factored out for reuse across the notebook and the UI.

Hugging Face distilBART Streamlit
View on GitHub →

Plus a long tail of smaller ML and annotation work — spam detection, sentiment annotation, submarine life simulation, classic-algorithm notebooks. Browse all repositories →

// 03 — experience

Where I've worked

Data Analyst Intern

Jan 2025 – Aug 2025

Sudhir Mehrotra & Associates, Chartered Accountants · Bareilly, India (Hybrid)

  • Built Python ETL pipelines that automated financial workflows, cutting manual processing effort by 30%.
  • Developed a time-series cash-flow forecasting module that improved estimation accuracy by 15% over baseline.
  • Automated recurring Excel reporting with Python and VBA, eliminating multi-hour weekly manual tasks for the audit team.
  • Designed structured data-reporting systems for audit and compliance teams, improving traceability across reviews.
Python ETL Time-Series Forecasting Excel / VBA
// 04 — skills

Tools of the trade

🤖

LLMs & RAG

OpenAI GPT-4.1LLaMA (Groq)RAG pipelinesHybrid searchCross-encoder rerankChromaDBsentence-transformersPrompt engineeringCitation grounding
⚖️

LLM Eval & Reliability

LLM-as-a-judgeHallucination detectionFaithfulness / relevanceCitation metricsCI-gated thresholdsVersioned prompts
🔥

Deep Learning

PyTorchTensorFlowKerasHF TransformersLSTMAttentionMC DropoutDeep ensemblesUncertaintyXAI
👁️

Computer Vision & Robotics

OpenCVHSV color detectionWebots (e-puck)Visual SLAMPath planningClosed-loop control
</>

Languages

PythonSQLRBashJavaScript
📊

Data & Viz

pandasNumPyNetworkXMatplotlibSeabornJupyteripywidgetsGradioStreamlit
⚙️

MLOps & Tools

GitGitHub ActionsAWSLinuxHF SpacesREST APIsRAPIDS (GPU)
🗄️

Databases

MySQLPostgreSQLMongoDBChromaDB (vector)
// 05 — certifications

Verified credentials

AI Fluency: Framework & Foundations

Anthropic · Issued Jun 2026

  • The 4D framework for working effectively with AI
  • Delegation & task decomposition between human and model
  • Description: prompting, context, and iterative refinement
  • Discernment & diligence — evaluating and verifying AI output
View Certificate →

Claude 101

Anthropic · Issued Jun 2026

  • Claude model family & capability fundamentals
  • Effective prompting and context management
  • Projects, artifacts, and extended workflows
  • Practical applications and responsible use
View Certificate →

Fundamentals of Accelerated Data Science

NVIDIA · Issued Oct 2025

  • GPU-Accelerated Computing (RAPIDS Ecosystem)
  • Core Data Science Foundations
  • Applied Machine Learning
  • Accelerated Ecosystem Integration
View Certificate →

AWS Academy Cloud Operations

Amazon Web Services · Issued Nov 2022

  • Cloud infrastructure operations
  • Monitoring & management on AWS
  • Deployment & automation fundamentals
View Certificate →
// 06 — education

Academic background

Apr 2024 – May 2026

M.S. in Data Science

University of Arizona · Tucson, AZ

GPA: 3.78 / 4.00

Machine Learning · Applied NLP · Computational Linguistics · Data Mining & Discovery

2020 – 2024

B.Tech, Computer Science & Engineering (Software Engineering)

SRM Institute of Science and Technology · Chennai, India

CGPA: 8.51 / 10.00

// 07 — research

Research & publications

Under review · NeurIPS 2026 arXiv:2605.12898

When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method

S. H. Kilaru, S. T. Manikyala, R. Upadhyay, S. S. K. Ramavath, S. Nunavathu, D. Alharthi · May 2026

Building on homophily and structural balance theory, we formalize four LLM-based tie-formation mechanisms — sequential, global, local, and iterative — as distinct conditional distributions over edge sets. Using a fixed roster of 50 demographically grounded personas, we generate 192 verified directed networks across four cultural contexts, four prompt languages, three GPT-4.1 variants, and four prompting architectures.

  • Cultural framing measurably shifts inbreeding homophily and largest-component connectivity.
  • Political affiliation dominates tie formation under three of four methods; the global method substitutes age — prompt architecture behaves as a substantive sociological variable.
  • Model scale produces a stable divergence ranking (GPT-4.1 ↔ mini = 0.074, ↔ nano = 0.119), with the smallest variant differing qualitatively rather than just noisily.
GPT-4.1 NetworkX pandas Experimental Design
// 08 — résumé

Want the one-pager?

Download my latest résumé or open it in a new tab to view my experience, projects, and education.

Last updated: August 2026

// 09 — contact

Let's build something

Open to full-time AI / ML Engineer roles in the US (on-site, hybrid, or remote). If you're hiring — or just want to talk LLMs — reach out.