Still celebrating National IT Professionals Day with 3 months of free Premium Membership. Use Code ITDAY17

x
?
Solved

Getting HTML Source from Remote URL!

Posted on 1998-12-22
5
Medium Priority
?
190 Views
Last Modified: 2010-03-05
I'm trying to create a small search engine submission script.

How do I get the source of a remote URL? (and) How do I access this URL with Perl?

I want to access..say..www.altavista.com/submit.cgi?url=mypage.com, then search the returned HTML code for "Successful submission".

Thanks!

Magecast
0
Comment
Question by:magecast
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 3
5 Comments
 
LVL 5

Expert Comment

by:b2pi
ID: 1207107
Simplest case (Start with

perldoc lwpcook

if this is overly simple....)

#!/usr/bin/perl -w
use strict;

use LWP::Simple;

my($doc) = get('http://www.altavista.com/submit.cgi?url=mypage.com');
$/ = undef;
if ($doc =~ m/successful submission/i) {
    print "Hey, that worked!\n";
} else {
   print "Uhoh, need to get more complex\n";
}


0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207108
By the way, you now have 4 questions locked and waiting for your response.  It is polite to either accept a given answer, ask for further clarification, or reject the answer...
0
 

Author Comment

by:magecast
ID: 1207109
Ok, I tried installing libwww-perl-5.41 but I need URI and a bunch of other stuff (where can I get this?) and it was just way too complicated.

Is there a way to do this without bringing in all these other modules?

Or is there any easy way I can just get a non-dependant module that does this job?

Thanks!
Matt

Thanks for letting me know about the locked questions b2..didn't know I had any.  I'll grade em now.
0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207110
There is a way to do this without bringing in all those other modules, but you could also do it in assembly language, and not worry about any of the wheels that other people have invented.

If you're going to do any www work, you really want libwww and cgi::* installed.  

1.) Are you using activestate or the standard distribution?
2.) Do you have either visual c++ or borland c++
0
 

Accepted Solution

by:
colind earned 120 total points
ID: 1207111
This is stolen almost directly from "Perl 5 by Example".  It uses the Socket module, but that's it.

sub http_get{
($_) = $address;           # Usage: httpget URL
($site, $url) = /^http:\/\/([^\/]*)(\/.*)/;
$_ = $site; ($desthost, $port) = /([^:]*)(.*)/;
$port = 80 unless $port;        # Default http port is 80
use Socket;
#chop($thishost = `hostname`);
$proto = (getprotobyname('tcp'))[2];
$port = (getservbyname($port, 'tcp'))[2] unless $port =~ /^\d+$/;
$thisaddr = (gethostbyname($thishost))[4];
$thataddr = (gethostbyname($desthost))[4] ||
        return "Unknown host '$desthost'";
        $this = pack('S n a4 x8', AF_INET, 0, $thisaddr);
        $that = pack('S n a4 x8', AF_INET, $port, $thataddr);
        # Create the connection to a remote server.
        socket(S, PF_INET, SOCK_STREAM, $proto) || print "socket: $!";
        bind(S, $this) || print "bind: $!";
        connect(S, $that) || print "connect: $!";
        select(S); $| = 1; select(stdout);
        print S "GET $url\r\n"; # Send request to get the document.
        while(<S>) { print output;
        print;}   # Read back the result.

}

0

Featured Post

Free Tool: Port Scanner

Check which ports are open to the outside world. Helps make sure that your firewall rules are working as intended.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Email validation in proper way is  very important validation required in any web pages. This code is self explainable except that Regular Expression which I used for pattern matching. I originally published as a thread on my website : http://www…
I have been pestered over the years to produce and distribute regular data extracts, and often the request have explicitly requested the data be emailed as an Excel attachement; specifically Excel, as it appears: CSV files confuse (no Red or Green h…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…
Six Sigma Control Plans

722 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question