Writing
Notes on building clinical AI that has to survive contact with actual hospitals, plus the occasional post about getting into rooms I was never invited to.
-
The transcript did not predict any of it
Arshi Niyaz had average marks at a rural CBSE school, with her father's consistent support. That school transcript did not predict the AKTU honours degree that followed. On snapshots, slopes, remote research access, and a Stanford summer traineeship completed from India. She is Co-Founder and Chief Science Officer at Biotech Wallah.
8 min read -
Operations is where the science survives
Trial logistics is research design wearing a different hat. Three ways a study quietly stops being valid in the corridor: throughput pressure produces missingness that flatters you, recruitment fixes the confound structure before any feature exists, and protocol drift is silent and cumulative. Jamyang Ugyen Tshomo has joined Biotech Wallah as Chief Operating Officer, owning research and operations together.
7 min read -
Sovereign models need sovereign benchmarks
India is building its own models. The evaluation is still borrowed. Notes on Sarvam's published stack, the codemix bet inside Saaras v3, and what a clinical benchmark would have to measure to be worth anything.
7 min read -
Skin tone fairness starts at the camera, not the model
Fairness in dermatology AI gets treated as a loss function problem. Most of the failure happens before the model sees anything, in the dataset and in the phone camera pipeline itself.
6 min read -
The three minute constraint
None of the things that decide whether a screening tool gets used in a busy outpatient department appear in the paper. Field notes from a trial in a tertiary hospital in India.
6 min read -
Clinical notes do not translate
Annotating Indian, Malaysian and Igbo clinical notes for a healthcare language model dataset. The same illness is written in three different grammars, and a model trained on one reads the others confidently wrong.
6 min read -
A screen that is 95 percent accurate is wrong about half the time
Accuracy in an abstract tells you almost nothing about a screening tool. Work the base rate and a very good model still hands you a majority of false positives. The arithmetic, and what to report instead.
7 min read -
Focus is advice for people who have not found the overlap
I run four companies and everyone tells me to pick one. The advice is correct for most people and wrong for me, and there is a single test that tells you which case you are in.
6 min read -
The wrong person to tell you about my brother
My brother started building at eleven, left college at eighteen because he decided the syllabus was years behind what he already did daily, and now runs engineering at my company. I am his sibling, so discount me accordingly. Here is what I think actually explains it, and where I think he is wrong.
12 min read -
The booth is the cost
A hearing test is one of the cheapest procedures in medicine, so why does almost nobody in rural India get one? The answer is the room, not the test. What has to be true before you can delete the sound booth.
7 min read -
A clinician who cannot see why will not act
I had a screening model that scored better than the one we deployed. I threw it away. On buying attribution with accuracy, and why explainability is an architectural constraint rather than a research topic you add later.
6 min read -
ICD-11 has no clean map from ICD-10
Coding is where health data quality is silently won or lost, long before anybody trains a model on it. Four migration cases, three of which need a clinician, and why we kept a human in the loop on purpose.
7 min read -
How I cold-emailed my way into eleven research labs
No application portals, no referrals, no pedigree. The first fifty emails got me nothing. Here is the single change that took the reply rate to roughly one in six.
6 min read