Solved

Scanned pages into pdf

Posted on 2006-06-27
3
310 Views
Last Modified: 2010-04-17
hello all does anyone  know how are scanned pages containing text converted into pdf and what is the format in which text is stored in the pdf. Can this text be directly extracted from the pdf or some technique like OCR is required to be used.

Thanks
0
Comment
Question by:jhav1594
3 Comments
 
LVL 1

Accepted Solution

by:
jm021196 earned 500 total points
ID: 16994505
It reallly depends on what app you are using and the quality of the page.

If the PDF Converting program which is being used to take the image from the scanner can recognise the text as text then its stored as text in the PDF File.

If the converting program cannot recognise it as text then it gets saved in a variety of image formats depending on which one suits it best. There really is no way to tell how its saved in advance.

PDF Files use a combination of vector, raster and text formates to give the best compression and viewability and so converting to PDF is a very difficult thing to undo... especiall if its not possible to tell in advance if its going to be in text or not.

I would suggest that a OCR system is the best way forward.

Thanks
mitch
0

Featured Post

Is Your Active Directory as Secure as You Think?

More than 75% of all records are compromised because of the loss or theft of a privileged credential. Experts have been exploring Active Directory infrastructure to identify key threats and establish best practices for keeping data safe. Attend this month’s webinar to learn more.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Displaying an arrayList in a listView using the default adapter is rarely the best solution. To get full control of your display data, and to be able to refresh it after editing, requires the use of a custom adapter.
Whether you've completed a degree in computer sciences or you're a self-taught programmer, writing your first lines of code in the real world is always a challenge. Here are some of the most common pitfalls for new programmers.
An introduction to basic programming syntax in Java by creating a simple program. Viewers can follow the tutorial as they create their first class in Java. Definitions and explanations about each element are given to help prepare viewers for future …
In this seventh video of the Xpdf series, we discuss and demonstrate the PDFfonts utility, which lists all the fonts used in a PDF file. It does this via a command line interface, making it suitable for use in programs, scripts, batch files — any pl…

910 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

23 Experts available now in Live!

Get 1:1 Help Now