Turning a photograph of words back into words you can edit
Reading text off a picture is one of those jobs that is either almost perfect or oddly bad, and the difference is nearly always the picture rather than the settings. This article covers what happens to your image before it is read, what the tick box actually changes, how to judge the confidence figure, and how to photograph a page so the result needs almost no correcting. A full worked example runs through a printed receipt from start to finish.
The everyday jobs this solves
The pattern is always the same: words exist, but not in a form you can select and copy. A paper receipt whose figures need to go into a spreadsheet. A page of a book you want to quote accurately without typing it out. An error message in a screenshot somebody sent you, so you can search for it. The serial number on the back of a router, photographed because the label is in an awkward corner.
There is also the meeting-room case. A slide is on screen, nobody will send the deck, so you photograph it from your seat and lift the text out afterwards. Same for a printed notice, a menu or a timetable. In each case the goal is not a beautiful copy of the page. It is the words, in a box, ready to paste somewhere useful.
Getting a picture in
There are three ways, and all of them end in the same place. Drag an image file onto the box, click it and choose a file, or press Ctrl and V to paste a screenshot straight from the clipboard. The paste route is the fastest for anything already on your screen, because you never make a file at all.
Once the picture is accepted it appears in the panel, and a line tells you its pixel size and its weight on disk. Those numbers are worth a glance: a picture only 600 pixels wide is going to struggle, and now is the moment to find a bigger one.
Anything that is not an image is refused with a short message rather than a silent failure. Clear empties the picture, the text box and the status line so you can start again.
Clean up the picture first, the only setting on the page
This single tick box, on to begin with, does two things before the reading starts. It turns the picture grey, and it stretches the range between the lightest and darkest parts so they reach true white and true black.
That second part is the useful one. Photograph a page under normal room lighting and the paper is not white; it is light grey, and the print is dark grey. The gap between ink and paper is narrow, and a narrow gap is what makes letters hard to separate. Pushing the two apart helps more than any adjustment inside the reading engine.
Leave it on for anything taken with a camera. The one case for turning it off is a screenshot that is already crisp black text on plain white, where there is nothing left to stretch. If a result disappoints, running it again with the box in the other position takes seconds and is always worth trying.
What happens to your picture before it is read
Your image is not handed over as it is. It is first brought into a size range that suits reading, and the small grey line under the buttons tells you exactly what was sent, in pixels, and whether the clean-up was applied.
Pictures that are too small get enlarged. Text needs a certain height in pixels before letters can be told apart at all, and a phone screenshot of small print often falls under it, so anything below roughly 1200 pixels on its long edge is scaled up to reach that.
Very large pictures get reduced, down to about 2600 pixels on the long edge. That surprises people, but a 6000 pixel photo of a page brings no extra readable detail and costs a lot of time. This is why a careful 2000 pixel photo often beats a casual 12 megapixel one.
The first run, and every run after it
Nothing is fetched until you actually press Read the text. That is deliberate: you should not pay for a download you never asked for just by opening a page.
The first press collects the two models the reader needs, and a bar reports each stage as it goes. It happens once. Your browser keeps the files, so the second use starts immediately, and every use after that too. One model finds where the text sits on the picture; the second reads each piece it found, which is why the bar names two stages rather than one.
The pleasant side effect is that this page then works with no connection at all. Load it once, go offline, and it still reads your pictures on a train or a plane. Everything it needs comes from this site rather than borrowed from elsewhere, so no other party learns what you are reading.
A page of ordinary printed text takes a few seconds on a laptop. An older phone takes noticeably longer.
Reading the line above the text box
When it finishes, a short summary appears: how many words were found, a confidence percentage, and how long it took. The middle number is the one to look at, because it is the engine's own opinion of how sure it is.
Above roughly 85 per cent the text is usually reliable and a quick skim is enough. Between 60 and 85, expect a handful of wrong characters and read it properly. Below 60 a warning appears on screen as well, and that is a signal to take a better picture rather than to start correcting by hand.
If nothing is found at all, the page says so and suggests a sharper or larger picture, or ticking the clean-up box. An empty result is almost never a fault in the reading; it means the letters were too small, too soft or too pale to be letters.
Worked example: a paper receipt
A receipt from a hardware shop, 24 items, needs to go into a spreadsheet. Lay it flat, ideally on something dark so the edges are obvious, and photograph it from directly above with the phone held level. Move in so the receipt fills the frame.
Drop that photo on the box. The status line reads something like 3024x4032 and 2.6 MB, then asks you to press Read the text. Leave Clean up the picture first ticked, because this is a camera photograph of paper.
Press the button. The bar runs through its stages and the small line reports about 1950x2600 px sent to the engine, cleaned up first. A few seconds later the text appears with a summary reading roughly 128 words, 89 per cent confident, 4.2 seconds.
Now check the numbers, because that is where mistakes gather. Look for a zero read as the letter O, a one read as a lower case L, a five read as an S. Three prices need a character corrected. Click into the text box, fix them, then press Copy and paste the lot into your spreadsheet. Download as .txt saves it as a file instead, named extracted-text.txt.
The text box is yours to edit
The result is not locked. You can click in and correct it before copying, which is quicker than pasting bad text elsewhere and fixing it there. Add a missing heading, delete the address you do not need, join a line that broke oddly.
Line breaks from the original are kept, and runs of blank lines are collapsed so the text is usable without tidying. Column layouts are another matter, covered below.
Copy places everything on your clipboard. Download as .txt writes a plain text file any program can open. Neither button touches the picture, so you can flip the tick box, read again and compare two attempts.
Taking a picture that reads well
Accuracy is decided at the moment of the photograph far more than at any point afterwards. Five habits cover almost everything.
- Get close. The text should fill the frame. Distant text in a wide shot is the single most common cause of poor results.
- Hold the camera square to the page, not at an angle. Sloping lines of text read much worse than level ones.
- Flatten the page. A curved book spine bends the lines and the letters with them.
- Watch your own shadow. Side lighting from a window beats standing between the lamp and the page.
- Crop before you come here if the picture contains a lot besides the text, and rotate it so the lines run across rather than up.
What it cannot do, said plainly
Handwriting is the big one. This is built for printed and typed text. Careful block capitals sometimes come through; ordinary joined-up writing generally does not, and no setting changes that.
It reads text, not page design. Line breaks survive, but two columns of a newsletter, a table of figures, or a paragraph wrapped around a photograph often come out in a jumbled order. The fix is to crop one column at a time and read each separately.
PDF files cannot be opened here directly. Take a screenshot of the page you want and paste that instead. On languages, the reader carries English and the Western European languages, Greek, Chinese and Japanese, but it holds no Cyrillic or Hangul characters at all, so Russian and Korean come back as nonsense rather than as a poor attempt. Handwriting remains hard for any reader and results will be patchy.
Pages that pair well with this one
If your photo of a page is skewed, shadowed or shot at an angle, run it through the Document Scanner first. It finds the corners, straightens the perspective and evens out the lighting, and that often turns a poor read into a clean one.
The Crop step in Batch Image Editor is the companion for splitting a two-column page into two single columns, and it too works without anything leaving your device.
Why this one insists on running locally
Think about what people actually feed an image-to-text tool. Receipts, payslips, letters from a bank, appointment cards, forms with an address on them. That list is a fair summary of the documents you should never hand to a free website for the sake of convenience.
So the reading happens inside this tab, on your own processor. The picture is opened from your drive and never sent anywhere, and the text exists only in your browser until you copy or save it. If you want proof rather than a promise, load the page, turn off your network, and read a picture anyway. It still works, and nothing that works offline can be sending your documents to anyone.
