We help IT Professionals succeed at work.

Regex and extract information from webpage (scraping)

nepaluz
nepaluz asked
on
513 Views
Last Modified: 2012-05-10
Hi,
I am trying to scrape the ISIN Number from yahoo finance pages and need some help.

I have built the url that I need but can not figure out how to scrape this information. I have attached a text file (which essentially is the web page) and here is the link to the page.

http://uk.finance.yahoo.com/q?s=prty&m=L&d=
Comment
Watch Question

Commented:
download the html page, then
use something like ReadAllText function to read the html page's source, then
use a InStr function to find "ISIN" - the value your after is just infront of this location in the text string,
and is ended by a closing paretheses - which you may easily find too

hope this makes sense to you, quite some work is yet to be done...good luck
you could use xpath to access the text for the ISIN:

/html/body/div/div[2]/div[4]/div[3]/div/div/span

If you really need a regex for it then here you go:
^.*ISIN (.{12}) \).*$

Author

Commented:
ErezMor thanks for the contrib. I will give it a bash andsee what I come up with.

smash_pants (what a name!)I really am interested in the regex but can not get it to work. Heres the code I am using (and don't laugh!)


Dim fxk = Regex.Matches(sr.ReadToEnd.ToString, "^.*ISIN (.{12}) \).*$")

Open in new window

Author

Commented:
OK - I have been un-able to implement either way suggested here, and though they may be correct, I am un-able to award poits for them. I am thus closing this thread and will try another one.

Commented:
since the last post from user was "i'll give it a try and get back to you...", i dont think his closing reason is justified.
i happen to have some actual experience in doing just what i suggested the user, so in terms of if it works or not, there's no doubt here.
unless i didnt understand his question, or his closing request reason...

Erez.

Author

Commented:
OK then erez, I have been unable to use the instr as suggested (I think I have to repeat the obvious), could you post some code to this end?
I've tested this code and it works.
Sorry i haven't been around much... I'm moving house.

string result = null;
string url = "http://uk.finance.yahoo.com/q?s=prty&m=L&d=";
WebResponse response = null;
StreamReader reader = null;

try {
    HttpWebRequest request = (HttpWebRequest)WebRequest.Create(url);
    request.Method = "GET";
    response = request.GetResponse();
    reader = new StreamReader(response.GetResponseStream(), Encoding.UTF8);
    result = reader.ReadToEnd();
}
catch (Exception ex) {
    // handle error
    Console.WriteLine(ex.Message);
}
finally {
    if (reader != null)
        reader.Close();
    if (response != null)
        response.Close();
}
foreach (Match txt in Regex.Matches(result, @"^.*ISIN (.{12}) \).*$", RegexOptions.Multiline)) {
    Console.WriteLine(txt.Groups[1]);
}

Open in new window

Unlock this solution and get a sample of our free trial.
(No credit card required)
UNLOCK SOLUTION

Author

Commented:
Thanks a lot and happy moving! If I may ask for a tweak, is it possible to have the regex NOT limiting the result string to 12 characters as you have it, but pich up to the closing bracket?

The ISIN (could?) be longer than 12 characters long and it would thus return an incomplete numer (or I am wrong?)

I have awarded the marks anyhow.
in the regex:
^.*ISIN (.{12}) \).*$

just change the .{12} to what you need.

Capital letters an numbers at least 12 chars long:
[A-Z0-9]{12,}

any character ,any length:
.*


Gain unlimited access to on-demand training courses with an Experts Exchange subscription.

Get Access
Why Experts Exchange?

Experts Exchange always has the answer, or at the least points me in the correct direction! It is like having another employee that is extremely experienced.

Jim Murphy
Programmer at Smart IT Solutions

When asked, what has been your best career decision?

Deciding to stick with EE.

Mohamed Asif
Technical Department Head

Being involved with EE helped me to grow personally and professionally.

Carl Webster
CTP, Sr Infrastructure Consultant
Empower Your Career
Did You Know?

We've partnered with two important charities to provide clean water and computer science education to those who need it most. READ MORE

Ask ANY Question

Connect with Certified Experts to gain insight and support on specific technology challenges including:

  • Troubleshooting
  • Research
  • Professional Opinions
Unlock the solution to this question.
Thanks for using Experts Exchange.

Please provide your email to receive a sample view!

*This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

OR

Please enter a first name

Please enter a last name

8+ characters (letters, numbers, and a symbol)

By clicking, you agree to the Terms of Use and Privacy Policy.