You cannot select or copy the text in a scanned PDF
You try to select text in a document from the office scanner, and nothing happens. Searching finds no matches. Dragging selects a rectangle instead of words.
The reason is that the file is just an image of paper. A scanner saves what the page looks like. To your eye it is text; inside the file it is no different from a photograph. There is no character data at all, so there is nothing to select, copy or search.
Getting it out requires reading the characters out of the image. Faint printing, skew and low resolution all produce passages that cannot be read. PDFIntact marks those as unreadable rather than leaving them blank — filling them in silently is what makes mistakes impossible to notice.
Try it now
Check it on your own PDF
The check runs entirely inside your browser. Your file is never uploaded, and no account is needed. See how many pages, and how many tables and figures, can be extracted before you pay.
Example
まず見分けてください
同じ「PDF」でも、原因によって対処が変わります。次のどれに当てはまるかで判断できます。
- 文字を選択できない・検索しても出てこない
- スキャンPDF(画像)です。このページの対象です。
- 選択もコピーもできるが、貼り付けると読めない文字になる
- 画像ではなく、フォントの対応表が欠けている状態です。PDFをコピーすると文字化けする をご覧ください。
- 文字は取り出せるが、表の行と列がずれる
- PDFが表の構造を持っていないためです。PDFの表をExcelに貼り付けるとずれる をご覧ください。
Frequently asked questions
- Why can I not copy text from a scanned PDF?
- Because the PDF contains a photograph of the page rather than character data. It looks like text to you, but to the file it is a picture, so there is nothing to select or search.
- How do I make it selectable?
- The characters have to be recognised from the image. PDFIntact analyses each page and exports the text, tables and charts to Excel, Word, Markdown or JSON. You can check in your browser, before paying, whether your file can be read.
- Can you read handwriting?
- Documents made up entirely of handwriting are out of scope. Use it for documents that are mainly printed text. The check tells you before purchase what cannot be converted.
- Can tables in scanned documents be extracted?
- Yes, with rules and merged cells reproduced. Faint, skewed or low-resolution printing will leave passages unread; those are reported as unreadable rather than guessed.
- What happens to the file I send?
- The check runs entirely in your browser and the original is never uploaded. A PDF sent for conversion is deleted within 24 hours, and by design nothing reaches our servers until payment has completed.
Check any PDF against all symptoms/ What each format contains