Want to win a PS4? Go Premium and enter to win our High-Tech Treats giveaway. Enter to Win

x
?
Solved

perl parse webpage

Posted on 2007-12-06
5
Medium Priority
?
817 Views
Last Modified: 2008-02-01
i am trying to count the number of occurances of a the following sentence in the source code of a webpage.

<!-- changed logic:noMatch to logic:notEqual to prevent similar symbols from displaying together -->

my script looks like this but it just runs and runs, i have to cancel it.


use LWP::Simple;
my $content = get('http://www.website');
my $count= 0;
$count++ while $content =~/\s+changed logic:noMatch to logic:notEqual to prevent similar symbols from displaying together\s/;
  print $count;
0
Comment
Question by:mcgilljd
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 3
  • 2
5 Comments
 
LVL 39

Expert Comment

by:Adam314
ID: 20421644
How long do you give it?  It might be taking a while to get the webpage.

Add a few print statements so you can see where it hangs...  What output do you get from this?

use LWP::Simple;
$|=1;
print "Getting page...\n";
my $content = get('http://www.website');
my $count= 0;
print "Got page: " . length($content) . " bytes\n";
 
print "Checking for line...\n";
$count++ while $content =~/\s+changed logic:noMatch to logic:notEqual to prevent similar symbols from displaying together\s/;
print "count=$count\n";

Open in new window

0
 

Author Comment

by:mcgilljd
ID: 20421977
Getting page...
Got page: 2582373 bytes
Checking for line...

then it hangs

the line should occur about 2100 times
0
 
LVL 39

Accepted Solution

by:
Adam314 earned 2000 total points
ID: 20422605
You need a /g on the end of your regex:

$count++ while $content =~/\s+changed logic:noMatch to logic:notEqual to prevent similar symbols from displaying together\s/g;

Open in new window

0
 

Author Comment

by:mcgilljd
ID: 20422746
i changed code alittle to this.

Now it just counts infinetly if the search string is found.

Maybe i need to split content into lines?
print "Checking for line...\n";
while ($content =~/\s+changed logic:noMatch to logic:notEqual to prevent similar symbols from displaying together\s/)
		{
		$count++ ;
			
		print "count=$count\n";
		}					

Open in new window

0
 
LVL 39

Expert Comment

by:Adam314
ID: 20422888
The /g should cause it to count properly.  If the message is on multiple lines, you might need /s also.
0

Featured Post

Ask an Anonymous Question!

Don't feel intimidated by what you don't know. Ask your question anonymously. It's easy! Learn more and upgrade.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Many time we need to work with multiple files all together. If its windows system then we can use some GUI based editor to accomplish our task. But what if you are on putty or have only CLI(Command Line Interface) as an option to  edit your files. I…
Checking the Alert Log in AWS RDS Oracle can be a pain through their user interface.  I made a script to download the Alert Log, look for errors, and email me the trace files.  In this article I'll describe what I did and share my script.
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…
Six Sigma Control Plans

604 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question