- Designed the evaluation framework for a production Computer Use Agent — run telemetry, failure taxonomy, and error attribution driving prompt and architecture iteration; 62% autonomous task completion across 800+ real-world sites.
- Built a benchmarking framework comparing Gemini against the incumbent Claude-based agent on task accuracy and latency, driving model selection for enterprise deployments.
- Found the dominant failure category conflated environment limits with recoverable agent errors; recategorizing separated true blockers from fixable behavior and redirected improvement effort.
- Architected an end-to-end Sales Ops automation integrating a CUA with event-driven Power Automate pipelines triggered by incoming emails; cut deal processing time by 30 min/sale (~180 min/week).
Rohit Upadhya
LinkedIn |Github |Portfolio |+91 96154 22513 |rohitupadhya18@gmail.com
Work Experience
- Fine-tuned embedding models for domain-specific semantic retrieval, improving recall@1 from ~40% to 85%+ over off-the-shelf embeddings.
- Architected and deployed a containerized Azure RAG pipeline (Azure OpenAI, Azure AI Search, Cosmos DB, Docker, GitHub Actions) with dynamic query routing, continuously ingesting Confluence pages to serve low-latency, context-aware answers to enterprise users.
- Designed a semantic matching pipeline using cross-encoder rerankers, improving client match rate by 20% over baseline and replacing manual matching, thereby eliminating 8+ hours of weekly manual work.
- Developed a quiz generator app leveraging PEFT fine-tuned LLaMA and Mistral models, reducing manual workload by over 80%.
- Designed a relation extraction module using GPT-4 & fine-tuned LLaMA models, with a Streamlit UI for testing and demos; structured 12,000+ support interactions to improve response accuracy and speed.
- Architected and engineered an automated tutorial generator using YOLO-based UI detection, Azure TTS/STT, OpenAI APIs, streamlining content creation for onboarding and training.
- Led development of REST APIs to automate CMDB updates and incident workflows, improving operational efficiency across 30+ engineering teams.
- Served as SME for Clarity PPM across 100+ projects, streamlining governance and ensuring consistent project reporting.
- Replaced premium analytics tools with ElasticSearch/Kibana, saving $15K+ quarterly and enhancing ticket trend visibility.
Publications
Introduced LexCLiPR, a multilingual dataset + retrieval framework enabling paragraph-level cross-lingual retrieval of legally relevant ECtHR cases.
Image-processing and ML-based system for automated detection of diabetic retinopathy.
Education
- Thesis: Cross-Lingual Query based Paragraph Retrieval for European Court of Human Rights Judgements
- Courses: Machine Learning | Deep Learning | NLP | Legal NLP | Deep Generative Models | Uncertainty Quantification
- Thesis: Diabetic Retinopathy Detection
Projects
- Extended LexCLiPR to address a limitation of fine-grained legal retrieval: chunking judgments into paragraphs strips the document context that determines relevance.
- Modeled each judgment as its own paragraph-level graph, training dual 3-layer Graph Attention Networks — local 5-hop adjacency and global BERTopic topic nodes — under contrastive loss.
- Outperformed a fully fine-tuned bi-encoder using a frozen mDPR encoder: Recall@5% of 53.5 vs 44.4 on seen queries and 39.0 vs 30.5 on unseen, against a 22.6 zero-shot baseline.
- Showed combining local and global views beats either alone (53.5 vs 51.2 / 52.2), while per-layer fusion collapses to 35.2 through over-smoothing.
- Implemented a Llama-style decoder-only transformer from scratch — custom BPE tokenizer, RoPE, RMSNorm, SwiGLU, multi-head attention with causal masking — pretrained on FineWeb-Edu on a single NVIDIA A30.
- Diagnosed and fixed a training bug where EOS tokens were never appended at document boundaries, then ran continued pretraining to recover sequence termination behavior.
- Instruction-tuned on databricks-dolly-15k and released base and instruct checkpoints publicly.
- Measured WikiText-2 perplexity of 74.57 on the base checkpoint.
Skills
Languages: Python, JavaScript, SQL, Shell
LLMs & Agents: RAG, Computer Use Agents (CUA), Agent Evaluation & Benchmarking, PEFT/LoRA Fine-tuning, Instruction Tuning, Prompt Engineering, Cross-Encoders, Contrastive Learning
ML Frameworks: PyTorch, HuggingFace (Transformers, PEFT, Datasets), PyTorch Geometric
Retrieval & Data: FAISS, Azure AI Search, Elasticsearch, Weaviate, Cosmos DB, Kafka
Infra & DevOps: Azure (ML, Functions, OpenAI), Docker, GitHub Actions, CI/CD, Linux, AWS S3