[{"data":1,"prerenderedAt":1188},["ShallowReactive",2],{"article:scanning-bhatkhande-notation-with-claude-vision":3,"\u002Fblog\u002Fscanning-bhatkhande-notation-with-claude-vision-surround":1177},{"id":4,"title":5,"archived":6,"author":7,"body":8,"date":1162,"description":1163,"extension":1164,"image":1165,"meta":1166,"minRead":987,"navigation":1167,"path":1168,"published":1167,"seo":1169,"sitemap":1170,"stem":1171,"tags":1172,"__hash__":1176},"blog\u002Fblog\u002Fscanning-bhatkhande-notation-with-claude-vision.md","Scanning Bhatkhande Notation with Claude Vision",false,"Baljeet Singh",{"type":9,"value":10,"toc":1148},"minimark",[11,15,32,35,42,48,53,56,59,98,101,105,108,111,114,121,125,128,135,138,141,148,152,155,158,165,169,172,175,178,184,275,278,393,396,402,406,409,412,415,481,488,598,609,615,741,744,748,751,754,757,760,764,774,777,802,805,817,824,995,1003,1006,1015,1021,1024,1028,1031,1034,1041,1044,1050,1054,1057,1090,1093,1096,1100,1105,1108,1115,1119,1122,1128,1134,1144],[12,13,14],"p",{},"Hindustani classical and Sikh devotional music is written in Bhatkhande notation — Carnatic music uses a different system entirely. There are thousands of published books in it. None of it is machine readable, so anyone who wants to play from a book and edit what they play has to retype it by hand, mark by mark.",[12,16,17,18,25,26,31],{},"I run two subscription products where people write and play this notation — ",[19,20,24],"a",{"href":21,"rel":22},"https:\u002F\u002Fkirtannotation.com",[23],"nofollow","Kirtan Notation"," and ",[19,27,30],{"href":28,"rel":29},"https:\u002F\u002Fsangeetnotation.com",[23],"Sangeet Notation",", two brands on one codebase — so \"point your phone at the page and get editable notation back\" is the feature that would save subscribers the most time. It took a while to get working, and the thing that finally worked was not the thing I expected.",[12,33,34],{},"This is what comes out the other end — the same notation, structured, editable and playable:",[12,36,37],{},[38,39],"img",{"alt":40,"src":41},"Structured notation: swaras in Gurmukhi with octave dots above and below, meend arcs under grouped notes, lyrics beneath each beat, vibhag markers along the bottom, and a per-section taal — Teentaal here, Ektaal in the next section","\u002Fimg\u002Fnotation-grid.jpg",[12,43,44],{},[45,46,47],"em",{},"This isn't live for subscribers yet — only admin accounts can run a scan today. Wider release may come later.",[49,50,52],"h2",{"id":51},"why-ocr-does-not-apply","Why OCR does not apply",[12,54,55],{},"Bhatkhande isn't a character set with a Unicode block you can recognise.",[12,57,58],{},"The swaras are ordinary letters — Gurmukhi, Devanagari or Latin. If that were all, this would be a solved problem. But the meaning lives in what surrounds them:",[60,61,62,74,80,86,92],"ul",{},[63,64,65,69,70,73],"li",{},[66,67,68],"strong",{},"Octave dots"," above or below the letter, with different positions for mandra and taar, and a separate ",[45,71,72],{},"ati"," form beyond each",[63,75,76,79],{},[66,77,78],{},"Kan swars"," — grace notes, printed small and raised beside the main note",[63,81,82,85],{},[66,83,84],{},"Meend"," — a glide, drawn as a bar spanning several notes, which can cross a beat boundary",[63,87,88,91],{},[66,89,90],{},"Chhand"," markers above the beat",[63,93,94,97],{},[66,95,96],{},"Beat and row structure",", where the taal fixes how many beats a row holds",[12,99,100],{},"A dot above a letter isn't a glyph in a line of text. It's a spatial relationship, and so is nearly everything else that matters. The layout carries as much meaning as the letters do. OCR gives you the letters and throws away the music.",[49,102,104],{"id":103},"first-attempt-hand-the-model-the-photo","First attempt: hand the model the photo",[12,106,107],{},"I started multi-provider — OpenAI, Gemini, Anthropic — on the theory that this was a hard vision task and the biggest available model would win.",[12,109,110],{},"The output was confidently wrong in a specific way. Notes in roughly the right sequence. Octave markers dropped or invented. Kan swars silently absorbed into the note beside them. Beat counts that didn't match the taal. It was never gibberish, which made it worse — you had to know the notation to see that it was wrong.",[12,112,113],{},"Describing the rules in prose didn't help much. I wrote increasingly precise English about where dots sit and what a kan swar looks like. The model got more confident and not more correct.",[12,115,116,117,120],{},"That's the signal, in hindsight. When a model is wrong in a ",[45,118,119],{},"plausible, pattern-completing"," way rather than a confused way, it is usually filling a gap rather than misreading your instructions.",[49,122,124],{"id":123},"the-thing-that-worked-a-reference-in-the-same-modality","The thing that worked: a reference in the same modality",[12,126,127],{},"The model has never seen Bhatkhande notation. No amount of describing it in words changes that, because the problem is visual and the description isn't.",[12,129,130,131,134],{},"So I stopped describing and started showing. I generate a ",[66,132,133],{},"reference sheet from the notation font"," — every swara, every octave form, kan, meend, ghaseet, chhand, each rendered as it actually appears on a page — and send that image in the prompt alongside the photograph.",[12,136,137],{},"The model now has a visual lookup table for marks it was never trained on. It isn't learning the notation system. It's being handed a key, in the same modality as the problem, and doing the matching it is already good at.",[12,139,140],{},"This is the single highest-leverage change in the whole pipeline. It beat every prose description I wrote, it is inspectable when it fails — you can look at the reference sheet and see what was ambiguous — and regenerating it takes an afternoon rather than a fine-tuning run.",[12,142,143,144,147],{},"The general form, which I now reach for first: ",[66,145,146],{},"if the problem is visual, give the model a reference that is also visual."," Describing a domain in words is a substitute for knowledge, and a poor one.",[49,149,151],{"id":150},"domain-conditioning-make-wrong-answers-detectable","Domain conditioning: make wrong answers detectable",[12,153,154],{},"The second change narrows what a legal answer can be.",[12,156,157],{},"A raga admits certain notes and forbids others. A taal fixes the number of beats in a row. Both are known before the scan runs — the user picks them in the UI — so both go into the prompt as constraints, and both come back out as post-processing checks.",[12,159,160,161,164],{},"This does two things. The model produces fewer illegal answers. And the illegal ones it does produce become ",[45,162,163],{},"detectable",", because a note outside the raga or a row with the wrong beat count is mechanically wrong rather than a matter of judgement. You can't correct what you cannot detect.",[49,166,168],{"id":167},"where-it-runs","Where It Runs",[12,170,171],{},"All of this is a Supabase edge function — Deno, server-side, one HTTP endpoint the app calls.",[12,173,174],{},"The reason is mundane and non-negotiable: the Anthropic API key. A vision call from the browser means shipping the key to the browser, so the pipeline has to live somewhere the user cannot read. Everything else about the design follows from being in a function rather than in the client.",[12,176,177],{},"Two constraints came with it, and both are worth knowing before you copy this.",[12,179,180,183],{},[66,181,182],{},"A size ceiling."," The function rejects anything over 5MB and accepts only JPEG, PNG and WebP:",[185,186,191],"pre",{"className":187,"code":188,"language":189,"meta":190,"style":190},"language-ts shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","const ALLOWED_MIME_TYPES = ['image\u002Fjpeg', 'image\u002Fpng', 'image\u002Fwebp'];\nconst MAX_IMAGE_SIZE_BYTES = 5 * 1024 * 1024;\n","ts","",[192,193,194,249],"code",{"__ignoreMap":190},[195,196,199,203,207,211,214,217,221,223,226,229,232,234,236,238,241,243,246],"span",{"class":197,"line":198},"line",1,[195,200,202],{"class":201},"spNyl","const",[195,204,206],{"class":205},"sTEyZ"," ALLOWED_MIME_TYPES ",[195,208,210],{"class":209},"sMK4o","=",[195,212,213],{"class":205}," [",[195,215,216],{"class":209},"'",[195,218,220],{"class":219},"sfazB","image\u002Fjpeg",[195,222,216],{"class":209},[195,224,225],{"class":209},",",[195,227,228],{"class":209}," '",[195,230,231],{"class":219},"image\u002Fpng",[195,233,216],{"class":209},[195,235,225],{"class":209},[195,237,228],{"class":209},[195,239,240],{"class":219},"image\u002Fwebp",[195,242,216],{"class":209},[195,244,245],{"class":205},"]",[195,247,248],{"class":209},";\n",[195,250,252,254,257,259,263,266,269,271,273],{"class":197,"line":251},2,[195,253,202],{"class":201},[195,255,256],{"class":205}," MAX_IMAGE_SIZE_BYTES ",[195,258,210],{"class":209},[195,260,262],{"class":261},"sbssI"," 5",[195,264,265],{"class":209}," *",[195,267,268],{"class":261}," 1024",[195,270,265],{"class":209},[195,272,268],{"class":261},[195,274,248],{"class":209},[12,276,277],{},"Modern phone photos blow straight through that. So the client compresses before uploading, above a threshold, capping the long edge at 2048px:",[185,279,281],{"className":187,"code":280,"language":189,"meta":190,"style":190},"if (file.size > COMPRESS_ABOVE_BYTES) {\n  imageFile = await imageCompression(file, {\n    maxSizeMB: 1.5,\n    maxWidthOrHeight: 2048,\n    useWebWorker: true,\n  });\n}\n",[192,282,283,307,334,349,362,376,387],{"__ignoreMap":190},[195,284,285,289,292,295,298,301,304],{"class":197,"line":198},[195,286,288],{"class":287},"s7zQu","if",[195,290,291],{"class":205}," (file",[195,293,294],{"class":209},".",[195,296,297],{"class":205},"size ",[195,299,300],{"class":209},">",[195,302,303],{"class":205}," COMPRESS_ABOVE_BYTES) ",[195,305,306],{"class":209},"{\n",[195,308,309,312,315,318,322,326,329,331],{"class":197,"line":251},[195,310,311],{"class":205},"  imageFile",[195,313,314],{"class":209}," =",[195,316,317],{"class":287}," await",[195,319,321],{"class":320},"s2Zo4"," imageCompression",[195,323,325],{"class":324},"swJcz","(",[195,327,328],{"class":205},"file",[195,330,225],{"class":209},[195,332,333],{"class":209}," {\n",[195,335,337,340,343,346],{"class":197,"line":336},3,[195,338,339],{"class":324},"    maxSizeMB",[195,341,342],{"class":209},":",[195,344,345],{"class":261}," 1.5",[195,347,348],{"class":209},",\n",[195,350,352,355,357,360],{"class":197,"line":351},4,[195,353,354],{"class":324},"    maxWidthOrHeight",[195,356,342],{"class":209},[195,358,359],{"class":261}," 2048",[195,361,348],{"class":209},[195,363,365,368,370,374],{"class":197,"line":364},5,[195,366,367],{"class":324},"    useWebWorker",[195,369,342],{"class":209},[195,371,373],{"class":372},"sfNiH"," true",[195,375,348],{"class":209},[195,377,379,382,385],{"class":197,"line":378},6,[195,380,381],{"class":209},"  }",[195,383,384],{"class":324},")",[195,386,248],{"class":209},[195,388,390],{"class":197,"line":389},7,[195,391,392],{"class":209},"}\n",[12,394,395],{},"2048px is the number I would tune first if accuracy ever regressed. Compress a page of dense notation too hard and the octave dots — the smallest marks, one or two pixels across at the wrong scale — start disappearing into JPEG artefacts. The whole pipeline can be perfect and still fail because a dot did not survive the upload.",[12,397,398,401],{},[66,399,400],{},"Everything is base64 over JSON."," The image and the font reference both go to the model as base64, so the page is base64 twice — once into the function, once into the API call. It works, and it is the reason the size ceiling matters more than it first appears.",[49,403,405],{"id":404},"the-app-already-knows-the-grid","The App Already Knows the Grid",[12,407,408],{},"This is the one I would have missed if I had started from the model instead of from the product.",[12,410,411],{},"Scanning does not happen on a blank page. By the time someone uploads a photo they have already created the composition in the editor — picked the taal, named the sections, maybe typed the first few beats. The app knows the shape of what is on that page before it ever calls the model.",[12,413,414],{},"So it sends that shape:",[185,416,418],{"className":187,"code":417,"language":189,"meta":190,"style":190},"interface SectionHint {\n  name: string;       \u002F\u002F \"Sthayi\", \"Antara\"\n  rows: number;       \u002F\u002F how many rows this section occupies\n  emptyBeats: number; \u002F\u002F leading beats with nothing in them yet\n}\n",[192,419,420,431,448,463,477],{"__ignoreMap":190},[195,421,422,425,429],{"class":197,"line":198},[195,423,424],{"class":201},"interface",[195,426,428],{"class":427},"sBMFI"," SectionHint",[195,430,333],{"class":209},[195,432,433,436,438,441,444],{"class":197,"line":251},[195,434,435],{"class":324},"  name",[195,437,342],{"class":209},[195,439,440],{"class":427}," string",[195,442,443],{"class":209},";",[195,445,447],{"class":446},"sHwdD","       \u002F\u002F \"Sthayi\", \"Antara\"\n",[195,449,450,453,455,458,460],{"class":197,"line":336},[195,451,452],{"class":324},"  rows",[195,454,342],{"class":209},[195,456,457],{"class":427}," number",[195,459,443],{"class":209},[195,461,462],{"class":446},"       \u002F\u002F how many rows this section occupies\n",[195,464,465,468,470,472,474],{"class":197,"line":351},[195,466,467],{"class":324},"  emptyBeats",[195,469,342],{"class":209},[195,471,457],{"class":427},[195,473,443],{"class":209},[195,475,476],{"class":446}," \u002F\u002F leading beats with nothing in them yet\n",[195,478,479],{"class":197,"line":364},[195,480,392],{"class":209},[12,482,483,484,487],{},"None of it is guessed. ",[192,485,486],{},"rows"," falls straight out of the taal:",[185,489,491],{"className":187,"code":490,"language":189,"meta":190,"style":190},"const beatsPerRow = taal?.beats;\nif (!beatsPerRow) throw new Error(`no taal for section ${section.name}`);\nconst rows = Math.ceil(section.beats.length \u002F beatsPerRow);\n",[192,492,493,513,561],{"__ignoreMap":190},[195,494,495,497,500,502,505,508,511],{"class":197,"line":198},[195,496,202],{"class":201},[195,498,499],{"class":205}," beatsPerRow ",[195,501,210],{"class":209},[195,503,504],{"class":205}," taal",[195,506,507],{"class":209},"?.",[195,509,510],{"class":205},"beats",[195,512,248],{"class":209},[195,514,515,517,520,523,526,529,532,535,537,540,543,546,549,551,554,557,559],{"class":197,"line":251},[195,516,288],{"class":287},[195,518,519],{"class":205}," (",[195,521,522],{"class":209},"!",[195,524,525],{"class":205},"beatsPerRow) ",[195,527,528],{"class":287},"throw",[195,530,531],{"class":209}," new",[195,533,534],{"class":320}," Error",[195,536,325],{"class":205},[195,538,539],{"class":209},"`",[195,541,542],{"class":219},"no taal for section ",[195,544,545],{"class":209},"${",[195,547,548],{"class":205},"section",[195,550,294],{"class":209},[195,552,553],{"class":205},"name",[195,555,556],{"class":209},"}`",[195,558,384],{"class":205},[195,560,248],{"class":209},[195,562,563,565,568,570,573,575,578,581,583,585,587,590,593,596],{"class":197,"line":336},[195,564,202],{"class":201},[195,566,567],{"class":205}," rows ",[195,569,210],{"class":209},[195,571,572],{"class":205}," Math",[195,574,294],{"class":209},[195,576,577],{"class":320},"ceil",[195,579,580],{"class":205},"(section",[195,582,294],{"class":209},[195,584,510],{"class":205},[195,586,294],{"class":209},[195,588,589],{"class":205},"length ",[195,591,592],{"class":209},"\u002F",[195,594,595],{"class":205}," beatsPerRow)",[195,597,248],{"class":209},[12,599,600,601,604,605,608],{},"The taal fixes beats per row — sixteen for teentaal, eight for keherwa — so the number of sections plus the beat count per section ",[45,602,603],{},"is"," the grid. The pipeline refuses to run without a ",[192,606,607],{},"taalId"," for exactly this reason: without beats per row there is no grid to describe, and the model is back to inferring layout from pixels.",[12,610,611,614],{},[192,612,613],{},"emptyBeats"," is the small one I like most. It counts the leading beats that already have notes, so a half-typed section resumes rather than restarting:",[185,616,618],{"className":187,"code":617,"language":189,"meta":190,"style":190},"let emptyBeats = 0;\nfor (const beat of section.beats) {\n  if (beat.notes && beat.notes.trim()) break;\n  emptyBeats++;\n}\n\u002F\u002F If every beat is empty, reset to 0 — the whole section needs scanning.\nif (emptyBeats === section.beats.length) emptyBeats = 0;\n",[192,619,620,635,660,698,705,709,714],{"__ignoreMap":190},[195,621,622,625,628,630,633],{"class":197,"line":198},[195,623,624],{"class":201},"let",[195,626,627],{"class":205}," emptyBeats ",[195,629,210],{"class":209},[195,631,632],{"class":261}," 0",[195,634,248],{"class":209},[195,636,637,640,642,644,647,650,653,655,658],{"class":197,"line":251},[195,638,639],{"class":287},"for",[195,641,519],{"class":205},[195,643,202],{"class":201},[195,645,646],{"class":205}," beat ",[195,648,649],{"class":209},"of",[195,651,652],{"class":205}," section",[195,654,294],{"class":209},[195,656,657],{"class":205},"beats) ",[195,659,306],{"class":209},[195,661,662,665,667,670,672,675,678,681,683,685,687,690,693,696],{"class":197,"line":336},[195,663,664],{"class":287},"  if",[195,666,519],{"class":324},[195,668,669],{"class":205},"beat",[195,671,294],{"class":209},[195,673,674],{"class":205},"notes",[195,676,677],{"class":209}," &&",[195,679,680],{"class":205}," beat",[195,682,294],{"class":209},[195,684,674],{"class":205},[195,686,294],{"class":209},[195,688,689],{"class":320},"trim",[195,691,692],{"class":324},"()) ",[195,694,695],{"class":287},"break",[195,697,248],{"class":209},[195,699,700,702],{"class":197,"line":351},[195,701,467],{"class":205},[195,703,704],{"class":209},"++;\n",[195,706,707],{"class":197,"line":364},[195,708,392],{"class":209},[195,710,711],{"class":197,"line":378},[195,712,713],{"class":446},"\u002F\u002F If every beat is empty, reset to 0 — the whole section needs scanning.\n",[195,715,716,718,721,724,726,728,730,732,735,737,739],{"class":197,"line":389},[195,717,288],{"class":287},[195,719,720],{"class":205}," (emptyBeats ",[195,722,723],{"class":209},"===",[195,725,652],{"class":205},[195,727,294],{"class":209},[195,729,510],{"class":205},[195,731,294],{"class":209},[195,733,734],{"class":205},"length) emptyBeats ",[195,736,210],{"class":209},[195,738,632],{"class":261},[195,740,248],{"class":209},[12,742,743],{},"Same principle as the font reference sheet, arriving from a different direction. The model was guessing at layout. The layout was sitting in the editor the whole time.",[49,745,747],{"id":746},"fan-out-per-section-not-per-page","Fan Out Per Section, Not Per Page",[12,749,750],{},"Once you have the section list, the next decision makes itself.",[12,752,753],{},"One call over a whole page loses fidelity. The model does a good job on the first section and drifts through the rest — attention spread over a full page of dense marks is worse than attention on one section of it.",[12,755,756],{},"So each section gets its own call, with its own hint, in parallel. Better output, and latency stays close to a single call because they run concurrently.",[12,758,759],{},"That decision creates the next problem.",[49,761,763],{"id":762},"the-cost-problem-and-where-the-caching-actually-goes","The cost problem, and where the caching actually goes",[12,765,766,767,770,771,773],{},"Every section call needs the same context: the system rules, the font reference sheet, and the notation image itself. Naively, ",[45,768,769],{},"N"," sections means re-sending all of that ",[45,772,769],{}," times. The font reference and the page photo are the two largest things in the request.",[12,775,776],{},"Prompt caching fixes it, but only if the breakpoints are in the right order. The rule is that a cached prefix is matched by exact prefix, so everything stable has to come before anything that varies. Three breakpoints, in this order:",[778,779,780,786,796],"ol",{},[63,781,782,785],{},[66,783,784],{},"The stable system rules",", before any hint text. Hints vary per scan; the rules don't.",[63,787,788,791,792,795],{},[66,789,790],{},"The font reference image."," It sits before the notation image deliberately — a ",[45,793,794],{},"different"," page with the same taal, raga and settings still reads this much of the prefix.",[63,797,798,801],{},[66,799,800],{},"The notation image",", the last block every section of this scan shares.",[12,803,804],{},"The section-specific instruction goes after all three, so it never becomes part of the cached prefix.",[12,806,807,808,810,811,813,814,294],{},"Then the part I did not expect to need. If you build that prefix and immediately fan out ",[45,809,769],{}," concurrent calls, all ",[45,812,769],{}," race a cold cache and each one pays the cache-write premium. That is ",[66,815,816],{},"strictly worse than not caching at all",[12,818,819,820,823],{},"The fix is a warm-up call before the fan-out, with ",[192,821,822],{},"max_tokens: 0"," — the input is processed, the cache is written, and it returns immediately with empty content and no output tokens billed:",[185,825,827],{"className":187,"code":826,"language":189,"meta":190,"style":190},"await postMessages(apiKey, {\n  model,\n  max_tokens: 0,\n  system: prefix.system,\n  messages: [{\n    role: 'user',\n    \u002F\u002F Placeholder sits AFTER the last breakpoint, so it never becomes part of\n    \u002F\u002F the cached prefix the real section calls read back.\n    content: [...prefix.content, { type: 'text', text: 'warmup' }],\n  }],\n});\n",[192,828,829,844,851,862,879,890,906,911,917,976,985],{"__ignoreMap":190},[195,830,831,834,837,840,842],{"class":197,"line":198},[195,832,833],{"class":287},"await",[195,835,836],{"class":320}," postMessages",[195,838,839],{"class":205},"(apiKey",[195,841,225],{"class":209},[195,843,333],{"class":209},[195,845,846,849],{"class":197,"line":251},[195,847,848],{"class":205},"  model",[195,850,348],{"class":209},[195,852,853,856,858,860],{"class":197,"line":336},[195,854,855],{"class":324},"  max_tokens",[195,857,342],{"class":209},[195,859,632],{"class":261},[195,861,348],{"class":209},[195,863,864,867,869,872,874,877],{"class":197,"line":351},[195,865,866],{"class":324},"  system",[195,868,342],{"class":209},[195,870,871],{"class":205}," prefix",[195,873,294],{"class":209},[195,875,876],{"class":205},"system",[195,878,348],{"class":209},[195,880,881,884,886,888],{"class":197,"line":364},[195,882,883],{"class":324},"  messages",[195,885,342],{"class":209},[195,887,213],{"class":205},[195,889,306],{"class":209},[195,891,892,895,897,899,902,904],{"class":197,"line":378},[195,893,894],{"class":324},"    role",[195,896,342],{"class":209},[195,898,228],{"class":209},[195,900,901],{"class":219},"user",[195,903,216],{"class":209},[195,905,348],{"class":209},[195,907,908],{"class":197,"line":389},[195,909,910],{"class":446},"    \u002F\u002F Placeholder sits AFTER the last breakpoint, so it never becomes part of\n",[195,912,914],{"class":197,"line":913},8,[195,915,916],{"class":446},"    \u002F\u002F the cached prefix the real section calls read back.\n",[195,918,920,923,925,927,930,933,935,938,940,943,946,948,950,953,955,957,960,962,964,967,969,972,974],{"class":197,"line":919},9,[195,921,922],{"class":324},"    content",[195,924,342],{"class":209},[195,926,213],{"class":205},[195,928,929],{"class":209},"...",[195,931,932],{"class":205},"prefix",[195,934,294],{"class":209},[195,936,937],{"class":205},"content",[195,939,225],{"class":209},[195,941,942],{"class":209}," {",[195,944,945],{"class":324}," type",[195,947,342],{"class":209},[195,949,228],{"class":209},[195,951,952],{"class":219},"text",[195,954,216],{"class":209},[195,956,225],{"class":209},[195,958,959],{"class":324}," text",[195,961,342],{"class":209},[195,963,228],{"class":209},[195,965,966],{"class":219},"warmup",[195,968,216],{"class":209},[195,970,971],{"class":209}," }",[195,973,245],{"class":205},[195,975,348],{"class":209},[195,977,979,981,983],{"class":197,"line":978},10,[195,980,381],{"class":209},[195,982,245],{"class":205},[195,984,348],{"class":209},[195,986,988,991,993],{"class":197,"line":987},11,[195,989,990],{"class":209},"}",[195,992,384],{"class":205},[195,994,248],{"class":209},[12,996,997,999,1000,1002],{},[45,998,769],{}," cache writes become one write plus ",[45,1001,769],{}," reads.",[12,1004,1005],{},"Two details around it that are easy to get wrong:",[12,1007,1008,1011,1012,294],{},[66,1009,1010],{},"Do not warm a single-section scan."," One section shares its prefix with nobody, so the warm-up is pure overhead. The code checks ",[192,1013,1014],{},"sectionHints.length > 1",[12,1016,1017,1020],{},[66,1018,1019],{},"Build the prefix once."," Rebuilding it per section would produce equal strings, so it would work — but the cache depends on the bytes being identical, and building it once makes that invariant explicit instead of accidental.",[12,1022,1023],{},"A failed warm-up is swallowed and the scan proceeds cold. A cold cache is a cost regression, not a broken feature, and I would rather bill myself than fail a user's scan.",[49,1025,1027],{"id":1026},"parse-leniently-validate-strictly","Parse leniently, validate strictly",[12,1029,1030],{},"This is the part where I had to unlearn something.",[12,1032,1033],{},"I shipped strict JSON parsing first, because that is what you are supposed to do. In production it threw away responses that were substantively correct — the right notes, the right octaves, the right beat counts — because of a trailing comma, or because the model wrote two paragraphs of analysis before the JSON.",[12,1035,1036,1037,1040],{},"So the parser is deliberately forgiving. It handles pure JSON, code-fenced JSON, and JSON embedded in prose. When there are several fenced blocks it takes the ",[66,1038,1039],{},"last"," one, because earlier fences are usually examples the model quoted back at itself. It falls back across the response shapes.",[12,1042,1043],{},"And then the meaning is validated hard: beats against the taal, notes against the raga, octave consistency across the row.",[12,1045,1046,1047],{},"A response that parses cleanly but claims a note the raga forbids is far more dangerous than one with a stray comma — and strict JSON parsing catches exactly the wrong one of those two. ",[66,1048,1049],{},"Validate the semantics you care about, not the serialisation.",[49,1051,1053],{"id":1052},"post-processing-that-knows-the-domain","Post-processing that knows the domain",[12,1055,1056],{},"The last layer fixes classes of error the model reliably makes.",[12,1058,1059,1060,1062,1063,1066,1067,1070,1071,1074,1075,1070,1078,1081,1082,1085,1086,1089],{},"My favourite, because of how quietly it fails: the ",[45,1061,72],{}," octaves — a full octave above or below — are ",[66,1064,1065],{},"one"," character in storage, ",[192,1068,1069],{},"U"," or ",[192,1072,1073],{},"L",", not a doubled ",[192,1076,1077],{},"uu",[192,1079,1080],{},"ll",". The tokenizer reads a single octave character, so ",[192,1083,1084],{},"suu"," renders on screen exactly like ",[192,1087,1088],{},"su"," and plays a whole octave off. It looks right and sounds wrong. The prompt asks for the single-character form; a vision model will still sometimes double the letter; so the doubled form is folded back after the fact rather than stored.",[12,1091,1092],{},"There is also octave-consistency repair — a middle-octave note sitting between two upper-octave neighbours gets promoted, because a single-note octave dip in the middle of a phrase is far more likely to be a missed dot than a real leap — and raga-constrained note correction, and a pickup-beat octave fix.",[12,1094,1095],{},"None of this is clever. All of it is the accumulated residue of looking at wrong output and asking what the model was systematically getting wrong, rather than what it got wrong once.",[49,1097,1099],{"id":1098},"then-i-measured-it","Then I Measured It",[12,1101,1102],{},[45,1103,1104],{},"Added a few weeks after this was first published.",[12,1106,1107],{},"For a long time the honest answer to \"how accurate is it?\" was \"roughly 90% on printed pages in manual testing\" — which has no denominator, no definition of correct and no interval. It stayed that way for months, including while I was writing everything above.",[12,1109,1110,1111,294],{},"I eventually built the eval. It found three bugs before it produced a single number, showed me that effort and model choice both move the result more than I expected, and that the same configuration scored 59.4%, 59.4% and 90.6% on one page. That is ",[19,1112,1114],{"href":1113},"\u002Fblog\u002Fwhat-happened-when-i-measured-my-vision-pipeline","its own post",[49,1116,1118],{"id":1117},"what-transfers","What transfers",[12,1120,1121],{},"Three things I would tell someone starting on a domain a model has never seen:",[12,1123,1124,1127],{},[66,1125,1126],{},"Give it a reference in the same modality as the problem."," This is the one worth leading with. Describing your domain in prose is a substitute for knowledge and it does not work nearly as well as a lookup table the model can look at.",[12,1129,1130,1133],{},[66,1131,1132],{},"Make wrong answers detectable."," Domain constraints are worth more as validators than as instructions.",[12,1135,1136,1139,1140,1143],{},[66,1137,1138],{},"Cache-breakpoint order is an architecture decision, not a flag."," Ordering your prompt by volatility, warming the prefix before you fan out, and knowing when ",[45,1141,1142],{},"not"," to warm it, is the difference between caching helping and caching costing you money.",[1145,1146,1147],"style",{},"html pre.shiki code .spNyl, html code.shiki .spNyl{--shiki-light:#9C3EDA;--shiki-default:#C792EA;--shiki-dark:#C792EA}html pre.shiki code .sTEyZ, html code.shiki .sTEyZ{--shiki-light:#90A4AE;--shiki-default:#EEFFFF;--shiki-dark:#BABED8}html pre.shiki code .sMK4o, html code.shiki .sMK4o{--shiki-light:#39ADB5;--shiki-default:#89DDFF;--shiki-dark:#89DDFF}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html pre.shiki code .sbssI, html code.shiki .sbssI{--shiki-light:#F76D47;--shiki-default:#F78C6C;--shiki-dark:#F78C6C}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .s7zQu, html code.shiki .s7zQu{--shiki-light:#39ADB5;--shiki-light-font-style:italic;--shiki-default:#89DDFF;--shiki-default-font-style:italic;--shiki-dark:#89DDFF;--shiki-dark-font-style:italic}html pre.shiki code .s2Zo4, html code.shiki .s2Zo4{--shiki-light:#6182B8;--shiki-default:#82AAFF;--shiki-dark:#82AAFF}html pre.shiki code .swJcz, html code.shiki .swJcz{--shiki-light:#E53935;--shiki-default:#F07178;--shiki-dark:#F07178}html pre.shiki code .sfNiH, html code.shiki .sfNiH{--shiki-light:#FF5370;--shiki-default:#FF9CAC;--shiki-dark:#FF9CAC}html pre.shiki code .sBMFI, html code.shiki .sBMFI{--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B}html pre.shiki code .sHwdD, html code.shiki .sHwdD{--shiki-light:#90A4AE;--shiki-light-font-style:italic;--shiki-default:#546E7A;--shiki-default-font-style:italic;--shiki-dark:#676E95;--shiki-dark-font-style:italic}",{"title":190,"searchDepth":251,"depth":251,"links":1149},[1150,1151,1152,1153,1154,1155,1156,1157,1158,1159,1160,1161],{"id":51,"depth":251,"text":52},{"id":103,"depth":251,"text":104},{"id":123,"depth":251,"text":124},{"id":150,"depth":251,"text":151},{"id":167,"depth":251,"text":168},{"id":404,"depth":251,"text":405},{"id":746,"depth":251,"text":747},{"id":762,"depth":251,"text":763},{"id":1026,"depth":251,"text":1027},{"id":1052,"depth":251,"text":1053},{"id":1098,"depth":251,"text":1099},{"id":1117,"depth":251,"text":1118},"2026-08-30","Bhatkhande notation isn't a character set, so OCR doesn't apply and no model has meaningful training data on it. Four things took this pipeline from confidently wrong to useful, and the one that mattered most was giving the model a reference in the same modality as the problem.","md","\u002Fimg\u002Fbhatkhande-claude-vision.jpg",{},true,"\u002Fblog\u002Fscanning-bhatkhande-notation-with-claude-vision",{"title":5,"description":1163},{"loc":1168},"blog\u002Fscanning-bhatkhande-notation-with-claude-vision",[1173,1174,1175],"AI","Anthropic Claude","Vision","fu9NFooGIR_zquUMCEbzfzJ7HaP1_57qQt68r2vbBYg",[1178,1183],{"title":1179,"path":1180,"stem":1181,"description":1182,"children":-1},"Up & Running with Vue.js 2.0 by creating a simple blog application","\u002Fblog\u002Frunning-vue-js-2-0-creating-simple-blog-application-709","blog\u002Frunning-vue-js-2-0-creating-simple-blog-application-709","In this article we will be Up & Running with Vue.js 2.0 by creating a simple blog application",{"title":1184,"path":1185,"stem":1186,"description":1187,"children":-1},"Build an app with Laravel5 (backend) and Angularjs (frontend) - Part 1","\u002Fblog\u002Fseries-build-an-app-with-laravel5-backend-and-angularjs-frontend-part-1-480","blog\u002Fseries-build-an-app-with-laravel5-backend-and-angularjs-frontend-part-1-480","In this part 1 of the series we will take a look at how to build the api using Laravel",1790075468803]