Your router does not know your evals
Model routing is an evaluation problem. A generic router cannot know where smaller models are good enough for your workload, but a router trained on your own judgments can.
ReadDiscipline under uncertainty.
रियाज़ · प्रणाली · स्वर
AI infrastructure, open-source systems, and Hindustani classical music.
I study and build the systems behind modern AI — especially the messy economics of serving LLMs in production. I'm also a Hindustani classical vocalist trained since childhood, which probably explains the obsession with structure, timing, and improvisation.
≈ USD 1.90
~68% utilization
ताल
illustrative · cost/1M output tokens
latency is tempo · telemetry is drone
Currently
A snapshot, not a CV. Updated when the work moves.
Research
An independent, concurrency-aware look at what inference actually costs — measured, not assumed.
Featured paper · arXiv
Token price is not serving cost.
Public LLM cost calculators often reduce serving economics to a static token price or assumed utilization. This work studies how request rate, concurrency, latency SLOs, hardware, model architecture, and quantization interact to change the real effective cost of self-hosted LLM inference.
Open source
A read-only observer that tells operators the truth about serving cost under their own traffic.
Objective live telemetry + effective cost-per-million-token meter for vLLM servers.
A read-only observer for running vLLM servers that ingests Prometheus metrics and surfaces live effective LLM serving cost against the operator's actual traffic.
Systems + Sound
Before AI, there was riyaz. The habits transfer more than people expect.
Sound
रियाज़
Before AI, there was riyaz.
I started training in Hindustani classical music around the age of five — years of riyaz before any code. It was never a hobby I dabbled in. It's the first thing I went deep enough on to actually master, and that depth is where the rest of how I think came from: patience, structure, attention, and the ability to stay with complexity for a long time.
Systems
Improvisation inside structure, repetition without boredom, deep listening for the thing that is slightly off. The same instincts run both columns.
Writing
Short, sharp pieces on inference economics, tools, and practice.
Model routing is an evaluation problem. A generic router cannot know where smaller models are good enough for your workload, but a router trained on your own judgments can.
ReadWhat making cover songs under unreasonable deadlines taught me about invisible effort, legible differentiation, and learning to lead the orchestra instead of playing every instrument.
ReadWhy I do the expensive version of things: the proof you can rebuild it is the only asset worth holding, the trick is controlling where you collect your dopamine, and slightly-above-average is the flywheel's ignition.
ReadArc
Not a résumé timeline — the same instinct showing up in different rooms.
Hindustani classical vocal training starts. Years of repetition, structure, and listening before any code.
Scrapers, ETL over SEC filings, and software systems — learning to model messy real-world data and ship.
Propael and other product experiments. Taste for the painful real problem, not the demo.
A concurrency-aware cost methodology and vllm-cost-meter — measuring what serving actually costs.
Discipline under uncertainty.
For LLM serving economics, open-source telemetry, research collaboration — or just to talk systems and sound.