on-device · nothing leaves this page

Photo to score

Upload a photo of printed sheet music and see what an on-device AI model actually detects on it, pixel by pixel — staff lines, noteheads, clefs, stems and rests — rendered as a colored overlay over your own photo. Everything runs locally in this tab: the models are downloaded once (~45 MB total, quantized) and every pass afterward runs entirely in your browser.

This is an honest first cut, not a finished transcription. The two passes below show you what the neural network actually sees, pixel by pixel. The "Transcribe a staff" pass turns that into a real downloadable MusicXML file — but only one staff at a time (no grand-staff/polyphony reduction), and every note comes out as a uniform quarter note (real rhythm — durations, beams, dots — isn't detected yet). Clef and key aren't detected either; you pick them, which also doubles as your chance to correct a misread. See omr-bridging-research.md and photo-to-score-research.md for the research behind this page and what's still open. Built on oemer's two segmentation models (MIT license).

Drop a photo of sheet music

jpg / png / anything your browser decodes — or click to choose

drag on the photo below to crop to just the page — background clutter (desk, hands, the edge of a book) can get misread as staff lines, and cropping also means the model spends its resolution budget on the music instead of the background

Staff & symbols

not run yet

Symbol detail

not run yet

Transcribe a staff

Run the Staff & symbols pass above first. Pick one detected staff, tell it which clef and key it's in (this pipeline doesn't classify those yet — a wrong pick is a one-click fix, not a fight with a model), and generate a monophonic transcription: every detected notehead on that staff, left to right, as a uniform quarter note. Rhythm (note durations, beams) isn't detected yet — see omr-bridging-research.md.

not run yet