Scan settings for clean, small PDFs
Two settings decide almost everything about a scanned PDF: how finely the page is sampled, and how many colours each sample gets. Both are chosen at the scanner, and neither can be recovered afterwards.
What resolution should I scan at?
300 dpi for anything you want to read or search. It is the resolution recognition engines are tuned for, and it puts roughly twenty pixels into the height of ordinary body type, which is about what an engine needs to tell an e from a c.
Below 200 dpi recognition falls apart on normal print, and no amount of later processing puts the detail back. Above 300 the returns arrive slowly: 400 to 600 dpi earns its size on small footnote type, faint carbon copies and anything you may need to enlarge, and past 600 you are mostly buying megabytes.
Black and white, grayscale or colour?
Grayscale is the safe default. Black and white throws away every shade at the moment of scanning, by deciding pixel by pixel which side of a threshold it falls on, and a page with faint strokes, pencil or a grey stamp loses those parts permanently.
Choose black and white only for crisp printed text where size matters more than anything else, and colour where colour carries information: signatures in blue ink, highlighter, coloured form guidance, photographs, or anything you may have to show is not a photocopy.
Why is my scanned PDF so large?
Because every page is a picture rather than text. A colour A4 page at 600 dpi is around a hundred megabytes before compression, a quarter of that at 300 dpi, and a third of that again in grayscale. Resolution and colour mode are multipliers, and they multiply once per page.
Compression brings that down at a price you choose. JPEG at a moderate quality is invisible on photographs and shows as soft edges around small text; the lossless modes keep the letterforms exactly and shrink far less. Scan sensibly and compress once at the end, rather than scanning coarsely to save space.
What if the scan is already bad?
Rescan, if you still have the paper. Nothing downstream adds detail that was never sampled, and ten minutes at the scanner beats an afternoon spent processing a 150 dpi grayscale file into something legible.
If the paper is gone, work in the right order: straighten and clean the pages, run recognition on the best version you have, then compress. Recognition reads the image, so compressing first feeds it artefacts, and the text layer it produces wants reading before you rely on it.
Where to go from here
JPG to PDF assembles scanned pages into one document, OCR PDF adds the searchable layer, and Compress PDF belongs at the end rather than the beginning.
What DPI and quality change goes deeper into resolution, what OCR can read covers the recognition side, and making a PDF smaller takes the size problem on its own.