Set the scanner correctly first
Scan at 300 DPI for ordinary body text. Going to 600 DPI helps only for very small print, footnotes and fine tables; below 200 DPI accuracy drops sharply.
Choose greyscale rather than pure black and white. One-bit scanning throws away the intermediate tones that help the model separate faint ink from paper texture.
Turn off aggressive "document enhancement" modes. They often thin strokes to make pages look crisp, which erases the thin parts of letters.
Handle the physical pages
Flatten staples and folds. A crease running through a line of text produces a shadow that reads as a stray character.
For bound books, press the spine flat or use a book scanner. Curved text near the gutter is the single most common failure in scanned book pages.
Scan double-sided documents in the correct order, and check the page sequence before running extraction so the output does not need reordering afterwards.
Check reading order, not just characters
Multi-column documents are where scanned extraction most often goes wrong in a way that is easy to miss: every word is correct, but the sentences are spliced across columns.
Read the first and last sentence of each output paragraph. If they do not follow each other logically, the layout analysis merged columns and you should process the columns as separate crops.
Headers, footers and page numbers usually appear inline in the output. Removing them with a repeated find-and-replace is quicker than deleting them page by page.
What to expect on an archive job
For a 200-page typed report scanned at 300 DPI, expect nearly clean text with occasional errors around hyphenated line breaks and in tables.
For carbon copies, faxes and mimeographed pages, expect substantially more correction. These are low-contrast originals and no setting fully compensates.