Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people, just like you, are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
Solved

Getting HTML Source from Remote URL!

Posted on 1998-12-22
5
186 Views
Last Modified: 2010-03-05
I'm trying to create a small search engine submission script.

How do I get the source of a remote URL? (and) How do I access this URL with Perl?

I want to access..say..www.altavista.com/submit.cgi?url=mypage.com, then search the returned HTML code for "Successful submission".

Thanks!

Magecast
0
Comment
Question by:magecast
  • 3
5 Comments
 
LVL 5

Expert Comment

by:b2pi
ID: 1207107
Simplest case (Start with

perldoc lwpcook

if this is overly simple....)

#!/usr/bin/perl -w
use strict;

use LWP::Simple;

my($doc) = get('http://www.altavista.com/submit.cgi?url=mypage.com');
$/ = undef;
if ($doc =~ m/successful submission/i) {
    print "Hey, that worked!\n";
} else {
   print "Uhoh, need to get more complex\n";
}


0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207108
By the way, you now have 4 questions locked and waiting for your response.  It is polite to either accept a given answer, ask for further clarification, or reject the answer...
0
 

Author Comment

by:magecast
ID: 1207109
Ok, I tried installing libwww-perl-5.41 but I need URI and a bunch of other stuff (where can I get this?) and it was just way too complicated.

Is there a way to do this without bringing in all these other modules?

Or is there any easy way I can just get a non-dependant module that does this job?

Thanks!
Matt

Thanks for letting me know about the locked questions b2..didn't know I had any.  I'll grade em now.
0
 
LVL 5

Expert Comment

by:b2pi
ID: 1207110
There is a way to do this without bringing in all those other modules, but you could also do it in assembly language, and not worry about any of the wheels that other people have invented.

If you're going to do any www work, you really want libwww and cgi::* installed.  

1.) Are you using activestate or the standard distribution?
2.) Do you have either visual c++ or borland c++
0
 

Accepted Solution

by:
colind earned 30 total points
ID: 1207111
This is stolen almost directly from "Perl 5 by Example".  It uses the Socket module, but that's it.

sub http_get{
($_) = $address;           # Usage: httpget URL
($site, $url) = /^http:\/\/([^\/]*)(\/.*)/;
$_ = $site; ($desthost, $port) = /([^:]*)(.*)/;
$port = 80 unless $port;        # Default http port is 80
use Socket;
#chop($thishost = `hostname`);
$proto = (getprotobyname('tcp'))[2];
$port = (getservbyname($port, 'tcp'))[2] unless $port =~ /^\d+$/;
$thisaddr = (gethostbyname($thishost))[4];
$thataddr = (gethostbyname($desthost))[4] ||
        return "Unknown host '$desthost'";
        $this = pack('S n a4 x8', AF_INET, 0, $thisaddr);
        $that = pack('S n a4 x8', AF_INET, $port, $thataddr);
        # Create the connection to a remote server.
        socket(S, PF_INET, SOCK_STREAM, $proto) || print "socket: $!";
        bind(S, $this) || print "bind: $!";
        connect(S, $that) || print "connect: $!";
        select(S); $| = 1; select(stdout);
        print S "GET $url\r\n"; # Send request to get the document.
        while(<S>) { print output;
        print;}   # Read back the result.

}

0

Featured Post

Announcing the Most Valuable Experts of 2016

MVEs are more concerned with the satisfaction of those they help than with the considerable points they can earn. They are the types of people you feel privileged to call colleagues. Join us in honoring this amazing group of Experts.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

In the distant past (last year) I hacked together a little toy that would allow a couple of Manager types to query, preview, and extract data from a number of MongoDB instances, to their tool of choice: Excel (http://dilbert.com/strips/comic/2007-08…
Checking the Alert Log in AWS RDS Oracle can be a pain through their user interface.  I made a script to download the Alert Log, look for errors, and email me the trace files.  In this article I'll describe what I did and share my script.
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…

840 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question