Solved

Getting HTML Source from Remote URL!

Posted on 1998-12-22
5
189 Views
Last Modified: 2010-03-05
I'm trying to create a small search engine submission script.

How do I get the source of a remote URL? (and) How do I access this URL with Perl?

I want to access..say..www.altavista.com/submit.cgi?url=mypage.com, then search the returned HTML code for "Successful submission".

Thanks!

Magecast
0
Comment
Question by:magecast
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 3
5 Comments
 
LVL 5

Expert Comment

by:b2pi
ID: 1207107
Simplest case (Start with

perldoc lwpcook

if this is overly simple....)

#!/usr/bin/perl -w
use strict;

use LWP::Simple;

my($doc) = get('http://www.altavista.com/submit.cgi?url=mypage.com');
$/ = undef;
if ($doc =~ m/successful submission/i) {
    print "Hey, that worked!\n";
} else {
   print "Uhoh, need to get more complex\n";
}


0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207108
By the way, you now have 4 questions locked and waiting for your response.  It is polite to either accept a given answer, ask for further clarification, or reject the answer...
0
 

Author Comment

by:magecast
ID: 1207109
Ok, I tried installing libwww-perl-5.41 but I need URI and a bunch of other stuff (where can I get this?) and it was just way too complicated.

Is there a way to do this without bringing in all these other modules?

Or is there any easy way I can just get a non-dependant module that does this job?

Thanks!
Matt

Thanks for letting me know about the locked questions b2..didn't know I had any.  I'll grade em now.
0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207110
There is a way to do this without bringing in all those other modules, but you could also do it in assembly language, and not worry about any of the wheels that other people have invented.

If you're going to do any www work, you really want libwww and cgi::* installed.  

1.) Are you using activestate or the standard distribution?
2.) Do you have either visual c++ or borland c++
0
 

Accepted Solution

by:
colind earned 30 total points
ID: 1207111
This is stolen almost directly from "Perl 5 by Example".  It uses the Socket module, but that's it.

sub http_get{
($_) = $address;           # Usage: httpget URL
($site, $url) = /^http:\/\/([^\/]*)(\/.*)/;
$_ = $site; ($desthost, $port) = /([^:]*)(.*)/;
$port = 80 unless $port;        # Default http port is 80
use Socket;
#chop($thishost = `hostname`);
$proto = (getprotobyname('tcp'))[2];
$port = (getservbyname($port, 'tcp'))[2] unless $port =~ /^\d+$/;
$thisaddr = (gethostbyname($thishost))[4];
$thataddr = (gethostbyname($desthost))[4] ||
        return "Unknown host '$desthost'";
        $this = pack('S n a4 x8', AF_INET, 0, $thisaddr);
        $that = pack('S n a4 x8', AF_INET, $port, $thataddr);
        # Create the connection to a remote server.
        socket(S, PF_INET, SOCK_STREAM, $proto) || print "socket: $!";
        bind(S, $this) || print "bind: $!";
        connect(S, $that) || print "connect: $!";
        select(S); $| = 1; select(stdout);
        print S "GET $url\r\n"; # Send request to get the document.
        while(<S>) { print output;
        print;}   # Read back the result.

}

0

Featured Post

Technology Partners: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

On Microsoft Windows, if  when you click or type the name of a .pl file, you get an error "is not recognized as an internal or external command, operable program or batch file", then this means you do not have the .pl file extension associated with …
Many time we need to work with multiple files all together. If its windows system then we can use some GUI based editor to accomplish our task. But what if you are on putty or have only CLI(Command Line Interface) as an option to  edit your files. I…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…
Six Sigma Control Plans

623 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question