Applied AI researcher

Hi :) I'm DhruviI teach small models to do big things
I study reward design in reinforcement learning to teach small language models to call tools correctly. At JioHotstar, I build applied AI for real users.
Qwen2.5-3B · 1,000 held-out examples · tool-call accuracy
NMIMS · 2026Dean's Student ListCertificate of Merit
Top 10% of students
Say hi 👋
About
What happens if I try?
Underlined links open more
These days, most of my work comes down to two things:
A strict reward taught a 3B model to call tools, on one T4 GPU.
Read the paper →Ask a sports question and it writes the SQL, across Pro Kabaddi in five languages and ball-by-ball T20.
See the project →An HR coach that answers from handbooks, and a sales agent that picks shows for a brand and writes the deck.
HR coach →Sales agent →Research
4 AI research papers
Papers I once barely understood slowly turned into papers I wrote. Here's what each one is about, in plain words.
Advancing SLM Tool-Use Capability using Reinforcement Learning
- Cheap small models need exact tool-call formatting.
- GRPO training: any extra text earns zero reward.
- A 3B model reached 71.1% tool-call accuracy on one T4 GPU.
Towards Actionable Fashion AI: A Holistic, End-to-End Style Recommendation Assistant
- Seasonal colours and body shape become 20-25 clothing picks.
- Fine-tuned DINOv2: 61% seasonal-colour accuracy versus Gemini 2.5 Flash at 43%.
- Designed to run on a phone.
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model
- Recommend books from movies a user liked, even with little history.
- Combine DistilBERT text features with FastText genre features.
- MAE (mean absolute error): 0.69 at top 20%, beating text-only, genre-only and TF-IDF baselines.
Advanced AI-Based Detection and Tracking System (ADTS) for Crime Prevention and Identification in Real-Time Surveillance
- Detect people, recognise faces and track them across CCTV cameras.
- YOLOv8, MTCNN and FaceNet: 83 frames per second.
- IDF1 (identity-tracking score): 89.8%.
Awards & roles
Beyond the papers
College was never just about academics for me.
- Research Head, IEEE Robotics & Automation Society, NMIMS
- Student Coordinator, Student Research Group, NMIMS
- Sub-Head, Google Developer Student Clubs, NMIMS (400+ students)
Work
Things I've built
A mix of things running in production and experiments I just wanted answers to.
Text-to-SQL sports chatbots
- Started from zero, with a lot of research before building
- 12 seasons of Pro Kabaddi in English, Hindi, Hinglish, Gujarati or Marathi, plus ball-by-ball T20
- 96% right on our 150-question Kabaddi test
Cloud cost dashboard
- ~62 projects across GCP and AWS in one view
- About 4 days of manual work down to about 5 minutes
- Senior management uses it now
HR AI Coach
- Answers questions from company handbooks
- Keeps HR record facts separate from anything the AI writes, so the two never mix
Sales Agent
- Takes a brand's brief and picks the 5 best-fit shows
- Writes the pitch deck around them
- The sales team uses it now
Promo effectiveness dashboard
- Shows how a promo is actually landing
- Pulls signals from YouTube, Instagram, X, Reddit, Google Trends and Hotstar
- Reads the comments for sentiment
Episode condensation
- Works from the dubbing script
- Trims recaps, ad-break gaps and pauses
- One episode went from 24:33 to 13:56
Audio separation pipeline
- Benchmarked SAM-Audio against HTDemucs on a real production file; BS-RoFormer won out
- Splits TV audio into 6 stems, so crowd noise, applause and camera clicks come out while the music stays
- Ran on masters from 4 channels
Music catalog downloads
- Crawls a music library's full genre tree and saves each track as WAV to cloud storage
- Playwright automation, with ProcessPool running workers in parallel and systemd keeping long bulk downloads going
- A download lock so workers never clash, and every file checked against its track title and catalog ID
Music copyright detection
- Tested ACRCloud, AudioShake and Gemini against cue sheets I checked by hand
- AudioShake did best, 80.8-90.1% on music-only audio
- Moved timestamps out of the LLM and into Python, which took Gemini to 64%
Cricket AI commentary
- Proof of concept: AI describing a cricket match between balls
- A strict data contract: it only says what the live stats and scorebug back up
- No generic filler
HR onboarding videos
- Tried 5 lip-sync models, picked LatentSync
- With ElevenLabs voices and n8n, makes personalized onboarding videos in batches
Personal branding agent
- Co-built an LLM platform on n8n that writes, schedules and publishes social posts
- I designed the agent workflow nodes that make each post personal
Personalized news recommender
- Java and Spring Boot microservices with service discovery and an API gateway
- Runs on Docker and Kubernetes
- Recommendations from TF-IDF and each user's preferences
Multi-camera CCTV tracking
- Real-time face tracking across multiple CCTV feeds
- YOLOv8, MTCNN and FaceNet
- Became the ADTS paper
Side projects
Built for fun
- Hacker News → tweets. Every day it reads the top 15 Hacker News stories, sums them up and drafts 10 tweets from them. I built it with CrewAI and Streamlit. GitHub ↗
- Salary negotiation bot. It predicts your salary with ML, and then you haggle with an LLM bot over it. After that, a second bot negotiates the same profile with a boss bot, so you can see who got the better deal. GitHub ↗
Writing
Explained simply
Never Give Up: why easy problems eat the RL budget
A deep dive into the Matthew Effect, and how adaptive sampling gives hard problems another chance.
A coin flip made this model better at math. Or did it?
Notes on "Spurious Rewards", a paper about what RL training really teaches a model.

The polite sentence that broke my model
How we taught small language models to call tools with GRPO, and why the strictest rule worked best.

Say hi