Dalvoy, 0 to 1: turning a test score into a diagnosis
Dalvoy is a test-prep product for serious government exam aspirants, and the product I am currently building from zero as founding designer. Every competitor in the category ships previous-year papers and mock tests. After talking to more than 500 students, we bet that the questions were never the product. The analysis after the test was. Dalvoy is past 4 lakh+ registered users.
In government exam prep, everyone sells questions. Nobody explains the mistakes.
Dalvoy serves aspirants preparing for UPSC and other competitive government exams in India. It is a multi-year commitment where a student is ranked against lakhs of peers for a few thousand seats, so study time is rationed and anything that wastes it gets deleted.
The product covers the preparation loop: MCQ tests on previous-year papers and mocks, a result-analysis engine, Mains answer evaluation, current affairs, a Learn section, a news surface, and Dalvoy Shorts. The category is fragmented but uniform in one way. Competitors ship the same mocks, return a score, and stop. None of them tell a student what is going wrong in the way they answer. That gap is the opening.
A student finishes a mock, sees 45 out of 100, and has no idea what to do next.
- They know they did badly, not why. A score is a verdict with no reason attached.
- They cannot name what traps them. The pattern exists across months of tests, but no single test surfaces it.
- They cannot see themselves against peers, which is the comparison that actually changes a study routine.
- They have no view of their own answering behaviour: time spent per question, answers switched, and how their thinking degraded from question 1 to question 100.
One designer, two engineers, and a launch date that did not move.
- No design team. Every screen across Android and iOS is mine, so the constraint was never taste or alignment. It was how much one person could reach before launch.
- Two engineers, split across features. Every module I designed competed for the same two people, and time was the binding constraint. There is visible design debt in the product because of it.
- Peer analysis needs peer volume. Percentile and topper comparisons only mean something once enough students have attempted the same questions, so the most differentiated part of the product got stronger only as the base grew.
500+ student conversations, and one reframe.
We built this from repeated conversations with more than 500 aspirants about what they do in the hours after a mock test. The honest answer, most of the time, was that they look at the score and change nothing, because nothing in front of them says what to change.
The insight. Students were not short of questions. They were short of a diagnosis. Every product in the category answers how did I do. None answer what is wrong with the way I answer. So the test is not the product. The result page is.
The test is not the product. The result page is.
Nine analysis modules on one screen, ordered by what a student reaches for first.
The design problem was not what to include. It was what a demoralised student, thirty seconds after a bad test, is willing to read, and in what order. User testing settled the first fold: score, accuracy and time, then rank analysis, then confidence against reality. Ranking sits second rather than buried because comparison is what makes a student act.
Everything below is the diagnosis: recurring weak concepts, strengths, speed against accuracy benchmarked to the average and to toppers, mistake types, exam pattern traps, and a recommendation into practice or lessons.
The trade-off I accepted. Density and a long scroll. A shorter result page would have been easier to read and would have made Dalvoy the same product as everything else in the category. The bet was that this audience will scroll a long way for a real diagnosis and will not scroll at all for a prettier score card.
Capturing how sure a student felt, without stealing time from a timed test.
Confidence against reality needs data an answer sheet does not hold: whether the student knew the answer or landed on it. Asking afterwards produces recalled guesses. Asking during risks seconds a student cannot spare.
What I designed. The moment an answer is selected, four chips appear on it: Sure, 50-50, Eliminated, Fluke. One tap and the row disappears, returning on the next answered question. It reads as part of answering rather than as a survey.
That single interaction powers the sharpest line on the result page. Instead of a generic accuracy figure, a student sees that toppers avoid fluke questions entirely, that they attempted eight, and that it cost them 5.33 marks in negatives. That is a specific, changeable behaviour.
Designing the surface for a machine that grades handwritten answers.
UPSC Mains is written by hand and marked subjectively, so an aspirant normally waits days for a mentor to look at an answer, and many never get reviewed at all. The evaluation engine already existed when I joined. I did not build the model. I designed the surface around it: how a student submits, what happens during the wait, and how a machine's judgment gets presented so a student trusts it.
Students practise on paper, so submission starts with a photo of the handwritten sheet. Evaluation takes time, so the flow does not pretend to be instant: submissions move into a history split into ongoing and completed, and a student can leave and come back to their scores.
Presenting the grade was the real problem. A single number from a model, on a handwritten answer about a career-defining exam, is very easy to disbelieve. Three decisions carry it:
- The score is decomposed, not delivered. Introduction, Body, Conclusion and Presentation each carry their own mark, so a 5 out of 10 is auditable. A student can see the content was fine and the presentation was not.
- The verdict is framed against peers. Next to the mark sits the question's difficulty and what the top scorer managed, so a 5 on a question where the best answer got 6 reads as a hard question rather than a bad student.
- Feedback lands on their own paper. Examiner markup renders over the uploaded sheet, mirroring the red pen every aspirant already trusts.
Students told us where their study time was going. It was going to reels.
Shorts started as my initiative. Students volunteered, unprompted, that they lose long stretches to short-form video on social apps, and said something similar built around their syllabus would keep them in the learning path instead. A market scan backed it up.
The risk I designed against. Short-form inside a serious study product can pull a student away from the deep work that moves a rank. The brief was a low-friction reason to return between study sessions, not a replacement for them.
What happened. Average session length across the app moved from roughly 15 minutes to 24. That is an app-wide before and after, not a controlled test, and other work shipped in the same window. I claim the feature, the case for building it, and the constraint that kept it pointed at the syllabus.
A paywall that never interrupts a student mid-test.
Dalvoy is freemium with Dalvoy PRO on top, at ₹199 a month or ₹999 a year. Every surface carries a free allowance rather than a locked door: a capped number of tests, Mains evaluations, news, Shorts and learning path content.
The rule I held. The gate sits at the boundary of an action, never inside one. It does not interrupt a timed test. This audience is preparing for an exam where concentration is the scarcest thing they have, and an upgrade prompt mid-paper does not just annoy a student, it corrupts the attempt and the data that attempt was going to produce.
The trade-off. Capping by surface means a free user reaches the paywall later, and some never reach it. We gave up a faster path to conversion for a free tier that is genuinely usable, betting that a student who has already built a habit on Dalvoy is worth more when they convert.
What shipped, what is still in build, and what shipped rougher than designed.
Live today. Tests on previous-year papers and mocks, the result-analysis engine with rank and peer comparison, confidence capture, Mains answer evaluation, current affairs, the Learn section, news with NewsFlash, progress, Dalvoy Shorts, and the subscription.
Designed, not yet released. The fatigue analysis, which reads how a student's accuracy and pace change from the first question to the last and shows where the crash happens. It is the module I most want in front of students, because it describes behaviour a student cannot observe in themselves during a timed paper.
Rougher than designed. We launched against a fixed date with two engineers, and there is visible design debt because of it. Being live and improving beat being polished and late for a product that needed real students generating real peer data, but it is a debt and I am working it down against user feedback.
What I can show, and what it means.
Session length is the number I trust most at this stage. In a category where students ration study time deliberately, a longer session is a student choosing to stay, and it is the closest available signal that the product is doing work they value.
What I am deliberately not claiming. Dalvoy has other numbers I am not publishing here, because they are company-level results in a window where acquisition, content and a growing team all moved together, and I cannot cleanly separate design's contribution from the rest. When those hold across more cycles and I can attribute them properly, they belong on this page. Until then they do not.
The diagnosis works. The prescription does not, yet.
What is not working. The recommendation module at the bottom of the result page is the weakest thing I have designed here. It correctly identifies what a student got wrong, then hands them a generic route into practice or lessons. The analysis above it is specific and the action below it is not, which breaks the promise the rest of the page makes.
What I want to build next. Shorts-backed remediation: each identified mistake resolving into a short, specific explanation of what went wrong, in the format students already come back for. That closes the loop the product currently leaves open.
What I would do differently. I would have shipped the confidence capture before the analysis modules that depend on it. The most differentiated insights on the result page run on that one in-test interaction, and building it first would have given us a longer runway of behavioural data.