Blog
Oct 8, 2026-6 MIN READ
Can Indian OCR Read Music Notation?

Can Indian OCR Read Music Notation?

I tested Bodhan's indic-ocr and Sarvam Vision against Claude Opus 5.5 on 27 pages of Bhatkhande notation. Both read Gurmukhi well and still lost by 35 points, because the music lives in marks that aren't letters.

By Baljeet Singh

My notation app turns a photo of Bhatkhande notation into editable, playable notes. Claude Opus 5.5 does the reading today, and since the last round of measurement and the fixes that followed, it gets about 99% of beats exactly right, octave included.

It also costs about 8 cents a page. Two Indian companies have recently shipped models built for Indian scripts: Bodhan's indic-ocr (open weights, ₹0.20 a page on their API) and Sarvam Vision. If either could read this notation, the scanner would get a lot cheaper. So I measured.

The setup

The same eval harness as before: 27 pages with a known right answer, scored beat by beat. A beat counts only if the swara, komal/tivra and octave are all right. 24 pages are screenshots of compositions in my own app; 3 are photos of a printed songbook.

OCR models return text, not notation, so each one's output went to a cheap model (Gemini Flash, text only, no image) whose only job was to turn it into the app's format. A page an OCR couldn't read counts as all wrong, because that's what a user would get.

Beats rightCost per page
Claude Opus 5.599.8%~$0.08
Bodhan indic-ocr64.6%~$0.024
Sarvam Vision60.8%~$0.022 + Sarvam

Opus was measured on 6 of the pages, and the others on all 27. Two runs per setup wherever it mattered.

The music is in the marks

Here is what the models have to read: a grid of beats, notes on top and lyrics below.

Both models read the Gurmukhi letters well. The trouble is that the letters are the easy half. A dot under a note means the lower octave. A line under it means komal. A dot above means the upper octave. These aren't characters, so an OCR model has no slot for them.

What they do instead is interesting. They read the marks as the nearest Gurmukhi vowel signs: the komal line becomes an aunkar (ੁ), the upper-octave dot a tippi (ੰ), and Bodhan sometimes turns the lower-octave dot into a halant (੍). That's usable if the next step knows the convention. But they drop the marks often, and inconsistently: the same note gets its dot in one row and loses it in the next, and on a repeated pair like ਸੰਸੰ the second note's mark goes missing.

Drop one dot and the note plays an octave too high. Every other part of the reading can be perfect and the music is still wrong.

Bodhan vs Sarvam

They fail in different ways.

Bodhan refuses dense pages. It couldn't transcribe 6 of the 27 pages at all, failing with "truncated" or "repetitive", every time I tried. Rows of repeated ऽ and – cells look like a model stuck in a loop, and its guard trips. On a fresh set of sharper screenshots it was worse: 13 of 24. On the pages it does read, though, its text gets to 84.8%.

Sarvam never refuses, but skips. It returned something for every page, yet on 6 pages it left out whole sections. On the three book photos it scored 42.5% to Bodhan's 85.8%. Its schema-guided Extract mode let me describe exactly how to write the marks; it ignored that and wrote vowel signs anyway.

Bodhan is the better of the two, and cheaper with a known price.

As a hint, they help

The other way to use an OCR is as a second opinion. I gave Gemini Flash the image and the OCR text:

Beats right
Gemini Flash alone81.6%
Flash + Bodhan's text86.1%
Flash + Sarvam's text85.4%

About five points, mostly from fixing which letter is in a cell. What's left is almost all komal marks: 7.7% of beats, against 1.2% wrong letters and 0.3% wrong octave.

Some ideas that didn't survive:

  • Adding Bodhan's text to Opus: 99.8% → 100% on 6 pages. One beat, for 2 more cents a page.
  • Accepting a beat when two cheap readings agree, and sending the rest to Opus: not one page came back with both readings fully agreeing, so every page would still go to Opus.
  • Using the raag: pages with a raag scored far higher, which looked like a lever. Re-running those pages with the raag removed moved them from 96.3% to 95.8%. They were just easier pages.

What would actually work

Claude stays. But one result kept me interested: on the three real book photos, Flash with Bodhan's hint reached 98.2%, with 4 of 6 runs perfect. Three pages from one book decides nothing, and real songbook photos from more books are the next thing to collect. It's the one place a cheap path came close.

The longer route is to train one. Bodhan's model is open, small (0.8B parameters for the reader), and licensed for fine-tuning. It already reads Gurmukhi swaras and already turns marks into something. And I have something most OCR projects don't: a renderer that can draw unlimited notation with the exact answer attached, so the training data costs nothing to label. A version trained on Bhatkhande could read a page for a fraction of a cent, and might even run offline on a phone.

That's a few weeks of work for a saving of 8 cents a scan, so it waits until there are enough scans, or until offline scanning matters enough on its own. Handwriting would be separate again: a model trained on clean renders learns print, not a teacher's notebook.

For now the lesson is the same one the last post taught me, from the other side. General-purpose models are good at letters. This notation is letters plus a layer of meaning in dots and lines that no OCR was built to see, and that layer is the music.

© 2019-2026 Baljeet Singh. All rights reserved.