A scanned document and a text document can look identical on screen and behave completely differently. Understanding which one you have explains most of the frustration people run into with PDFs: why the file is enormous, why searching finds nothing, why compression barely helps, and why drawing a black rectangle over a name does not hide it.
Try to select a line of text with your cursor. If you get a text selection, the document contains real text. If you get a rectangular selection box or nothing at all, it is an image. You can also zoom in a long way: text stays crisp at any magnification because it is drawn from font outlines, while a scan becomes visibly pixelated.
A scanned page is a photograph of a piece of paper. At 300dpi a Letter page is roughly two and a half thousand by three and a quarter thousand pixels. In colour that is a substantial image before compression and a few megabytes after. Multiply by the page count. A fifty-page scanned contract at a hundred and fifty megabytes is not a broken file; it is what fifty photographs weigh. Structural PDF compression will save you almost nothing, because there is no structural waste — the size is the image data.
Most people scan at defaults that are far higher than needed. For a text document that will be read and archived, 200dpi in greyscale is entirely adequate and produces files perhaps a fifth the size of 300dpi colour. Reserve 300dpi colour for documents where the colour carries information or where the result may be reproduced. Many scanners also offer a text or document mode that applies contrast curves making paper read as white, which both improves legibility and compresses better.
Optical character recognition analyses the image, recognises letter shapes, and writes an invisible text layer behind the picture. The page still looks like a scan but the words are selectable and searchable. This is a genuinely hard computational task and is why OCR is not something a browser tool does casually. Many scanners and scanner apps offer it at scan time, which is the easiest place to get it. Accuracy is good on clean printed text and poor on handwriting.
Drawing a black rectangle over sensitive text in a PDF editor hides it visually and changes nothing underneath. On a text-based PDF the words are still there and can be recovered by selecting the region, copying, and pasting into a text editor. This has caused real disclosures in litigation and journalism with some regularity. Proper redaction deletes the underlying content and requires a tool built to do it. On a scanned page there is no text layer to recover, so a black box over an image is genuinely opaque — unless the document has been OCRed, in which case the text layer is there and searchable.
Scan at modest settings with OCR enabled if the scanner offers it. Keep the original scan. If pages need extracting or reordering, do it at the page level rather than re-scanning. If the file must be smaller and the scan is already modest, the honest options are splitting it or accepting a visible quality reduction, not a compressor promising to do it losslessly.