Solved

Need powershell script to scan muutple pdf's for keywords

Posted on 2013-11-25
9
1,917 Views
Last Modified: 2013-11-25
Greeting Experts,

I am in need of a simple PowerShell script to scan a folder full of pdf’s (2000 +) for keywords in the text of each one… Does somebody have a script or point me into the direction where I can find one script to complete this task…
0
Comment
Question by:amstoots
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 3
  • 3
  • 2
  • +1
9 Comments
 
LVL 35

Expert Comment

by:Dan Craciun
ID: 39674821
Why do you need a powershell script?
You can achieve the same goal using Windows search or any other piece of software that can do text search.

FWIW, on Windows, I use Notepad++ to search for text in folders.

HTH,
Dan
0
 

Author Comment

by:amstoots
ID: 39674929
the documents I am trying to scan are pdf's  and using the Windows Search only scans for names of the documents.. not the text inside of the documents....  that is what i am trying to do....
0
 
LVL 35

Expert Comment

by:Dan Craciun
ID: 39674976
OK. Here's how you do search in files in Notepad++:
Search in files in Notepad  You actually can use Windows Search to find in files, but with Notepad++ you have access to regular expressions, if need arises.
You can get Notepad++ for free from here: http://notepad-plus-plus.org/

HTH,
Dan
0
Guide to Performance: Optimization & Monitoring

Nowadays, monitoring is a mixture of tools, systems, and codes—making it a very complex process. And with this complexity, comes variables for failure. Get DZone’s new Guide to Performance to learn how to proactively find these variables and solve them before a disruption occurs.

 
LVL 40

Expert Comment

by:footech
ID: 39675125
BTW, you can scan inside the .PDFs with Windows Search as long as you have the right iFilter.  For 64-bit systems, Adobe has their version 11.
http://www.adobe.com/support/downloads/detail.jsp?ftpID=5542
If you have a 32-bit system, the iFilter comes with Adobe Reader.
0
 
LVL 54

Expert Comment

by:Joe Winograd, EE MVE
ID: 39675221
Dan,
I just tried to search the contents of PDFs with the latest Notepad++ (6.5.1) and it doesn't work. The PDF files do have text...searches with Adobe Reader (and other search tools) find the text, but not NPP. Please try it on your end and let me know your results. Thanks, Joe
0
 

Author Comment

by:amstoots
ID: 39675254
I did try to use notepad ++ and was unsuccessfully when I tried to scan the list of pdf's . after doing little bit of digging , I found a article that shows how to scan using adobe reader  

URLhttp://www.ghacks.net/2011/04/02/how-to-search-multiple-pdf-documents-at-once/
0
 
LVL 35

Accepted Solution

by:
Dan Craciun earned 500 total points
ID: 39675300
My bad. Was under the impression that PDF's conform to some xml standard, so they are text files with pictures encoded as binary (something like emails).

Turns out I was wrong: PDF's are binary files and the text is not directly readable from a text editor.

I apologize, I was spreading misinformation.
0
 

Author Closing Comment

by:amstoots
ID: 39675311
Hey, you helped point me in the right direction.. thanks...
0
 
LVL 54

Expert Comment

by:Joe Winograd, EE MVE
ID: 39675315
amstoots,
Yes, Adobe Reader can do it, as can other PDF readers/viewers (such as Foxit Reader and PDF-XChange Viewer), as well as many search products, such as dtSearch and X1, as well as the built-in Windows Search 4 (included with Vista/W7/W8 and available as a free download for XP).

Dan,
Thanks for confirming. Would be a nice enhancement for NPP7. :)

Regards, Joe
0

Featured Post

Free Tool: IP Lookup

Get more info about an IP address or domain name, such as organization, abuse contacts and geolocation.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

A project that enables an administrator to perform actions within a user session context not just at the time of login but any time later on day(s) or week(s) later.
In previous parts of this Nano Server deployment series, we learned how to create, deploy and configure Nano Server as a Hyper-V host. In this part, we will look for a clustering option. We will create a Hyper-V cluster of 3 Nano Server host nodes w…
Learn the basics of lists in Python. Lists, as their name suggests, are a means for ordering and storing values. : Lists are declared using brackets; for example: t = [1, 2, 3]: Lists may contain a mix of data types; for example: t = ['string', 1, T…
The viewer will learn how to count occurrences of each item in an array.

752 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question