Skip to content
Vipul Ponugoti

Vipul Ponugoti

AI Engineer — LLM systems, RAG, agentic workflows, evaluation & safety.

Final-year B.Tech CSE (AI & ML), graduating 2027 · SDE Intern @ VibeLevel.ai

Production-focused AI engineer building RAG platforms, agentic and MCP-based systems, LLM evaluation harnesses, and self-hosted inference APIs, shipped through Docker and Kubernetes.

Amber figures were fetched from GitHub when this page was built (2026-09-08). The others come from the resume dated 2026-08-25.

Projects

Temporal UI Evidence Graphs

Jul 2026
Problem
UI failures are often temporal: a request succeeds but state never updates, a route transition silently fails, a layout breaks only after a viewport change. Screenshots and flat event traces hide the order of events, so language models diagnose them poorly.
Built
A browser-instrumentation and benchmark pipeline that converts console, network, route, DOM, layout, and accessibility events into typed temporal evidence graphs, plus an evaluation harness that rebuilds every result offline from stored artifacts without provider calls.
Result
  • 3 experiments across 50 controlled scenarios, 3 models, and 10 evidence conditions
  • 2,700 stored diagnoses; 10,643 immutable artifacts verified
  • 9/9 offline reproducibility checks pass
  • Temporal-graph method ranked first in 9 of 10 bug categories

TypeScript, Node.js, Playwright, React

Repository not yet public

Enterprise Agentic RAG Platform

Problem
Answer questions over PDFs, web pages, and structured data with retrieval that is precise, access-controlled, and auditable down to the citation.
Built
Hybrid BM25 + dense retrieval with cross-encoder reranking over pgvector. Spring Boot microservices for JWT auth, RBAC, document management, and chat sessions; Python services for adaptive chunking, embeddings, and context-aware generation. LangGraph workflows for multi-hop document comparison, iterative query refinement, and citation-level source verification, with confidence scoring on every response.
Result
  • 91% RAGAS faithfulness on the internal benchmark
  • 87% context precision@5

Spring Boot, Python, LangGraph, pgvector, LlamaIndex, Redis, Docker, React

Code not yet pushed

MCP-Powered AI Project & Community Assistant

Jun – Jul 2026
Problem
Hosted models are disconnected from a team's current repositories and private documents, and sending internal files to a provider raises privacy and cost concerns.
Built
A local-first assistant that routes each request through a LangGraph router to specialised workers: GitHub analysis over MCP tools, document question answering with local embeddings and FAISS, web research, and issue workflows. REST and SSE streaming endpoints, Ollama inference, Docker Compose deployment, configuration validation, and automated tests.
Result
  • Deterministic routing with explicit tool boundaries; GitHub write actions require confirmation
  • Backend, Streamlit UI, and Ollama start from one docker-compose file; Apache-2.0

Python, FastAPI, LangGraph, MCP, FAISS, Ollama, Streamlit, Docker

OS Anomaly Sentinel

Jul 2026
Problem
Fixed CPU and memory thresholds raise false alarms on busy machines and miss real anomalies on quiet ones. Every machine has a different normal.
Built
Process-level telemetry collection with psutil, statistical time-window features, an Isolation Forest baseline learned per machine without labels, and a 10-page Streamlit dashboard for live monitoring, training, detection, investigation, and report generation.
Result
  • Isolation Forest compared against LOF, One-Class SVM, and fixed thresholds on precision, recall, F1, and confusion matrices
  • No kernel modifications, labelled data, or process termination required

FinOps Autopilot

Jan 2026
Problem
Idle and oversized cloud resources waste spend, but teams hesitate to act because automated remediation in production is risky.
Built
An agentic cost-optimisation service that detects waste across EC2, EBS, RDS and more, assesses each resource with LangGraph and GPT-4, remediates behind production guards with automatic rollback, and routes approvals through Slack. Policy-as-code with OPA/Rego, predictive cost analytics, and AWS Lambda deployment.
Result
  • Human-in-the-loop approvals and rollback over Slack
  • 36 test files across detection, remediation, and policy modules

Python, LangGraph, AWS, Azure, GCP, Slack, OPA/Rego

RailFlow OS

Jun 2026
Problem
When a disruption cascades across a rail corridor, operators need impact analysis and ranked recovery options with a safe approval step, not a chatbot.
Built
A control-tower prototype that keeps corridor state, simulates disruption propagation, ranks recovery plans with transparent scoring, asks a human operator to approve, generates advisories, and stores an audit trail. The demo corridor runs five stations with 8 passenger and 2 freight movements.
Result
  • Human-in-the-loop by design: the system never controls infrastructure directly
  • Backend and frontend run from docker-compose; Apache-2.0

TypeScript, Next.js, Python, Docker Compose

Open source

GSSoC '26 contributor (GirlScript Foundation, India) since May 2026: ranked 15 of 47,951 (top 0.03%) with 90,524 points in the AI/Agents and Open Source tracks. Peak of 200 merged pull requests in a single week, across Node.js, Python/FastAPI, Rust, Dart/Flutter, and Solidity codebases. Leaderboard.

Most of these pull requests are fixes: authentication and authorization, race conditions and TOCTOU bugs, rate limiting, input validation and sanitisation, secrets handling, and CI.

Merged pull requests per upstream repository, from GitHub as of 2026-09-08.

  • RepositoryMerged PRsStars
  • KanishJebaMathewM/Truxify35123

    Freight logistics platform (Node.js, Flutter): server-side freight pricing, JWT and WebSocket authorization, HMAC webhook verification on the raw body, path-traversal protection on ML model endpoints, PromQL input sanitisation, a TOCTOU race fix in request deduplication, and a Vitest unit-test framework.

  • kalyan-1845/ai-code-reviewer20420

    Multi-agent LLM code review platform: atomic circuit-breaker HALF_OPEN admission, symlink TOCTOU fix via O_NOFOLLOW, blocking Redis KEYS replaced with SCAN, safe YAML schema parsing, secret-scrubber and redirect-cap hardening, non-root containers, removal of committed dev credentials, CI test-artifact repairs.

  • RatLoopz/sahidawa-india5679

    Medicine verification platform: retry-and-throw on stale alert deactivation, pg_cron monitor initialisation fix, SSRF bypass via IPv6 addresses, authorization-header redaction in webhook logs.

  • Ayushh-Sharmaa/NexaSphere5234

    Tier rate-limiter guards, RBAC permission recording, singleflight deduplication of concurrent cache misses, SQL-injection guard fix, socket origin enforcement in production.

  • Ixotic27/The-Leetcode-City3975

    Auth and ownership checks on push routes, CRON_SECRET-protected jobs, rate-limit race fix, atomic XP-grant RPC, removal of a destructive DELETE from a GET endpoint.

  • harshdwivediiiii/pathfinder-ai2218

    Zod validation of LLM output before database writes, atomic rate-limit operations, TOCTOU fix in rate checks.

  • JhaSourav07/commitpulse18162

    Atomic INCR rate limiter replacing a get-then-set race, quota-leak fix, correct cache status for cold requests, strict ISO date validation.

  • durdana3105/peer-learning1631

    Duplicate-review prevention and related correctness fixes.

  • viru0909-dev/nyay-setu-working1584

    Timing-safe JWT signature comparison, strict CSP headers, X-Forwarded-For rate-limiter spoofing fix.

  • Puneet04-tech/AegisGraph-Sentinel-2.01313

    torch.load with weights_only=True to close a pickle deserialisation RCE, removal of a hardcoded SUPER_ADMIN auth bypass, router registration, structured logging.

  • Kallappa2005/MLOPS_RED_WINE_QUALITY_PREDICTION87

    Quality gate on model promotion, canonical stable-model rollback, real MLflow failures surfaced instead of mock runs, corrected feature-drift ratio, hardened input validation.

  • param20h/PDF-Assistant-RAG750

    JWT leak in an export query parameter, IDOR on document update, SSRF DNS-rebinding bypass, WebSocket origin validation, pickle RCE in BM25 index deserialisation.

  • SdSarthak/AegisAI696

    SSRF prevention on webhook URLs, atomic bulk-scan commits, timezone-aware timestamps.

  • All repositories (44)895

Experience

SDE Intern, VibeLevel.ai

Jul 2026 – present, remote

AI-native skills and talent platform (practice assessments, vibeathons, hiring funnel): Next.js App Router on Vercel, FastAPI on Fly.io, Neon Postgres with Drizzle, E2B code sandboxes, and an MCP server. About 24 merged commits in a team codebase.

  • Designed and built Aura Scoring v2, an LLM-based human–AI collaboration scoring engine: a multi-phase pipeline (attribution analysis, human-contribution corroboration, substance ceiling, hypothesis provenance) that corrected agency misattribution in AI-assisted sessions, where v1 scored a passive delegator and a strategist identically. Shipped with about 2,700 lines of pytest coverage across 12+ test files.
  • Built knowledge-graph-grounded session feedback: extracted workspace context and grounded LLM scoring insights in a product feature graph via a custom PFG client.
  • Shipped a full-stack workspace signals feature aggregating agent-reported frameworks into user profiles, with GitHub repo derivation from session tags (FastAPI, Next.js, TypeScript, React).
  • Hardened the MCP server's contracts, evidence verification, and client compatibility; added a read-only scoring-service status endpoint and closed floating-point gaps in score-band lookup.
  • Virtualised the hiring candidate table with rAF-based row windowing (with tests) and fixed the auth layer to distinguish database outages from an unauthenticated state, eliminating spurious logouts.

open-aura on GitHub

ML Engineer Trainee Intern, Kshemam Health Solutions

Feb 2026 – Jun 2026, remote
  • Owned an end-to-end predictive risk classification pipeline on patient-monitoring data: privacy-compliant feature engineering, scikit-learn model evaluation, and precision/recall tuning for real-time clinical-risk thresholds.
  • Built Java/Spring Boot REST endpoints and scheduled batch pipelines for automated AI analytics; contributed to data-ingestion and model-serving workflows with attention to reproducibility, security, and inference latency.
  • Established model performance baselines, latency benchmarks, and structured error-analysis reports where no systematic tracking existed before.

Machine Learning Intern, Cognifyz Technologies

Dec 2025 – Jan 2026, remote
  • Applied data-analysis and supervised-learning tasks in Python with repeatable preprocessing, training, and evaluation workflows in pandas and scikit-learn.

More projects

  • AI Safety & Constitutional Evaluation Harness

    Red-teaming framework and constitutional filtering pipeline for adversarial prompt testing across multiple LLM providers. Classifies outputs on safety, helpfulness, and harmlessness axes and produces audit reports used to harden system prompts.

    Python, PyTorch, FastAPI, Docker

  • Self-Hosted LLM Inference & Deployment System

    vLLM-backed inference API on Kubernetes with streaming, batching, per-tenant rate limiting, and token metering. QLoRA fine-tuning of Llama-3 8B and Phi-4 on domain datasets, monitored with Prometheus and Grafana.

    vLLM, PyTorch, Kubernetes, LoRA/QLoRA, Prometheus, Grafana

  • LLM Evaluation & Observability Platform

    Automated evaluation harness over dataset, model, retriever, and chunking configurations, tracking RAGAS faithfulness, context recall, citation accuracy, hallucination rate, and p95 latency, with CI quality gates that block releases below configurable thresholds.

    Python, FastAPI, Spring Boot, PostgreSQL, LangSmith, Weights & Biases, Docker

  • Distributed Event-Driven Task & Workflow Engine

    in progress

    Kafka-backed distributed task scheduler with idempotency keys and transactional PostgreSQL job claims, dead-letter queues, exponential-backoff retries, circuit breakers, Redis token-bucket rate limiting, and OpenTelemetry/Prometheus telemetry.

    Spring Boot, Kafka, Redis, PostgreSQL, Docker, OpenTelemetry

  • Finlinter

    Static analysis that treats cloud cost as a code bug: detects cost-heavy patterns in Python, JavaScript, and Java, estimates cost, and suggests fixes before deployment.

    Python

  • SocialConnect Backend

    REST API for a social matching platform: profiles, preferences, interaction tracking, a paginated activity feed, S3 asset storage, and Flyway-managed schema evolution.

    Java 17, Spring Boot 3, Flyway, AWS S3

  • Distributor Hub

    Analytics app for retail distributors with demand forecasting, trend classification, and recommendation scoring served through ONNX Runtime.

    React, TypeScript, Express, ONNX Runtime

Skills

Languages
Python, Java, TypeScript, JavaScript, SQL
AI & LLM engineering
RAG (BM25 + dense hybrid, pgvector, FAISS, Pinecone, LlamaIndex, reranking), Agentic workflows (LangGraph, AutoGen), Model Context Protocol (MCP), Prompt engineering, RAGAS evaluation, Constitutional AI, RLHF concepts, Hallucination detection, Red-teaming
ML frameworks
PyTorch (primary), Hugging Face PEFT/Transformers, TensorFlow, scikit-learn, pandas, LoRA/QLoRA fine-tuning, Phi-4, Llama-3, Statistical modelling, Feature engineering
Backend & APIs
Java, Spring Boot, Spring AI, FastAPI, Flask, Node.js/Express, REST, WebSocket, SSE, JPA/JDBC, SQLAlchemy, PostgreSQL (Neon, Drizzle ORM), MongoDB, Redis, Apache Kafka
MLOps, infra & testing
Docker, Docker Compose, Kubernetes, Fly.io, GitHub Actions CI/CD, AWS (EC2, S3, Bedrock), Azure, Oracle Cloud Infrastructure, vLLM, Ollama, LangSmith, Weights & Biases, MLflow, OpenTelemetry, Prometheus, Grafana, pytest, Playwright, Vitest, Git/GitHub, Linux
Frontend
React, Next.js (App Router), TypeScript, Streamlit, Vercel

Credentials

Certifications

  • Machine Learning Specialization

    DeepLearning.AI and Stanford Online, via Coursera. Three courses: supervised learning, advanced learning algorithms, unsupervised learning and recommenders.

    2025-05-06Verify
  • Mathematics for Machine Learning and Data Science Specialization

    DeepLearning.AI, via Coursera. Linear algebra, calculus, probability and statistics for ML.

    2026-03-15Verify
  • AWS Academy Graduate — Cloud Foundations

    AWS Academy, via Credly. 20-hour training badge.

    2026-03-06Verify
  • OCI 2025 Generative AI Professional

    Oracle

    2025Verify
  • OCI 2025 Data Science Professional

    Oracle

    2025Verify
  • OCI 2025 AI Foundations Associate

    Oracle

    2025Verify

Programs and coursework

  • Quantum Fundamentals Program

    WISER, with Amaravati Quantum Valley and Qubitech. Four-week program equivalent to 1 academic credit. Certificate ID 7B053459.

    2025–26
  • WISER Summer Program 2026: What is the future of optimization? Classical + AI + Quantum

    WISER. Selected among the top 3,000 of 65,000 participants.

    2026-07-23
  • MIT 6.S191 Introduction to Deep Learning

    MIT

    Coursework

Hackathons

Participant: ISRO Bharatiya Antariksh Hackathon 2026 (Hack2Skill); Cognizant Technoverse Hackathon 2026; Deccan AI Experts Catalyst Hackathon.

Education

B.Tech, Computer Science & Engineering (AI & ML)

2023 – 2027

Parul Institute of Engineering and Technology, Vadodara. Dean's List 2024–25.

Coursework: Machine Learning, Deep Learning, NLP, Probability & Statistics, Data Structures & Algorithms, Database Systems, Distributed Systems, Operating Systems, Computer Networks.

Beyond code

  • Guinness World Records participant

    Largest mental arithmetic lesson, 3,189 participants, Chennai, 14 October 2018.

  • National Abacus Championship Gold Topper, twice

    34th and 37th National Abacus Competitions, Brainobrainfest 2017 and 2018, Chennai.