Solved

Need powershell script to scan muutple pdf's for keywords

Posted on 2013-11-25
9
1,712 Views
Last Modified: 2013-11-25
Greeting Experts,

I am in need of a simple PowerShell script to scan a folder full of pdf’s (2000 +) for keywords in the text of each one… Does somebody have a script or point me into the direction where I can find one script to complete this task…
0
Comment
Question by:amstoots
  • 3
  • 3
  • 2
  • +1
9 Comments
 
LVL 34

Expert Comment

by:Dan Craciun
ID: 39674821
Why do you need a powershell script?
You can achieve the same goal using Windows search or any other piece of software that can do text search.

FWIW, on Windows, I use Notepad++ to search for text in folders.

HTH,
Dan
0
 

Author Comment

by:amstoots
ID: 39674929
the documents I am trying to scan are pdf's  and using the Windows Search only scans for names of the documents.. not the text inside of the documents....  that is what i am trying to do....
0
 
LVL 34

Expert Comment

by:Dan Craciun
ID: 39674976
OK. Here's how you do search in files in Notepad++:
Search in files in Notepad  You actually can use Windows Search to find in files, but with Notepad++ you have access to regular expressions, if need arises.
You can get Notepad++ for free from here: http://notepad-plus-plus.org/

HTH,
Dan
0
3 Use Cases for Connected Systems

Our Dev teams are like yours. They’re continually cranking out code for new features/bugs fixes, testing, deploying, testing some more, responding to production monitoring events and more. It’s complex. So, we thought you’d like to see what’s working for us.

 
LVL 39

Expert Comment

by:footech
ID: 39675125
BTW, you can scan inside the .PDFs with Windows Search as long as you have the right iFilter.  For 64-bit systems, Adobe has their version 11.
http://www.adobe.com/support/downloads/detail.jsp?ftpID=5542
If you have a 32-bit system, the iFilter comes with Adobe Reader.
0
 
LVL 52

Expert Comment

by:Joe Winograd, EE MVE
ID: 39675221
Dan,
I just tried to search the contents of PDFs with the latest Notepad++ (6.5.1) and it doesn't work. The PDF files do have text...searches with Adobe Reader (and other search tools) find the text, but not NPP. Please try it on your end and let me know your results. Thanks, Joe
0
 

Author Comment

by:amstoots
ID: 39675254
I did try to use notepad ++ and was unsuccessfully when I tried to scan the list of pdf's . after doing little bit of digging , I found a article that shows how to scan using adobe reader  

URLhttp://www.ghacks.net/2011/04/02/how-to-search-multiple-pdf-documents-at-once/
0
 
LVL 34

Accepted Solution

by:
Dan Craciun earned 500 total points
ID: 39675300
My bad. Was under the impression that PDF's conform to some xml standard, so they are text files with pictures encoded as binary (something like emails).

Turns out I was wrong: PDF's are binary files and the text is not directly readable from a text editor.

I apologize, I was spreading misinformation.
0
 

Author Closing Comment

by:amstoots
ID: 39675311
Hey, you helped point me in the right direction.. thanks...
0
 
LVL 52

Expert Comment

by:Joe Winograd, EE MVE
ID: 39675315
amstoots,
Yes, Adobe Reader can do it, as can other PDF readers/viewers (such as Foxit Reader and PDF-XChange Viewer), as well as many search products, such as dtSearch and X1, as well as the built-in Windows Search 4 (included with Vista/W7/W8 and available as a free download for XP).

Dan,
Thanks for confirming. Would be a nice enhancement for NPP7. :)

Regards, Joe
0

Featured Post

Netscaler Common Configuration How To guides

If you use NetScaler you will want to see these guides. The NetScaler How To Guides show administrators how to get NetScaler up and configured by providing instructions for common scenarios and some not so common ones.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Set OWA language and time zone in Exchange for individuals, all users or per database.
This article explains how to prepare an HTML email signature template file containing dynamic placeholders for users' Azure AD data. Furthermore, it explains how to use this file to remotely set up a department-wide email signature policy in Office …
Learn how to match and substitute tagged data using PHP regular expressions. Demonstrated on Windows 7, but also applies to other operating systems. Demonstrated technique applies to PHP (all versions) and Firefox, but very similar techniques will w…
The viewer will learn how to count occurrences of each item in an array.

831 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question