Solved

Scanned pages into pdf

Posted on 2006-06-27
3
312 Views
Last Modified: 2010-04-17
hello all does anyone  know how are scanned pages containing text converted into pdf and what is the format in which text is stored in the pdf. Can this text be directly extracted from the pdf or some technique like OCR is required to be used.

Thanks
0
Comment
Question by:jhav1594
3 Comments
 
LVL 1

Accepted Solution

by:
jm021196 earned 500 total points
ID: 16994505
It reallly depends on what app you are using and the quality of the page.

If the PDF Converting program which is being used to take the image from the scanner can recognise the text as text then its stored as text in the PDF File.

If the converting program cannot recognise it as text then it gets saved in a variety of image formats depending on which one suits it best. There really is no way to tell how its saved in advance.

PDF Files use a combination of vector, raster and text formates to give the best compression and viewability and so converting to PDF is a very difficult thing to undo... especiall if its not possible to tell in advance if its going to be in text or not.

I would suggest that a OCR system is the best way forward.

Thanks
mitch
0

Featured Post

Networking for the Cloud Era

Join Microsoft and Riverbed for a discussion and demonstration of enhancements to SteelConnect:
-One-click orchestration and cloud connectivity in Azure environments
-Tight integration of SD-WAN and WAN optimization capabilities
-Scalability and resiliency equal to a data center

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

This article is meant to give a basic understanding of how to use R Sweave as a way to merge LaTeX and R code seamlessly into one presentable document.
A short article about problems I had with the new location API and permissions in Marshmallow
In this fourth video of the Xpdf series, we discuss and demonstrate the PDFinfo utility, which retrieves the contents of a PDF's Info Dictionary, as well as some other information, including the page count. We show how to isolate the page count in a…

860 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question