[2 days left] What’s wrong with your cloud strategy? Learn why multicloud solutions matter with Nimble Storage.Register Now

x
?
Solved

how to extract all the hyperlinks on this webpage

Posted on 2012-12-29
7
Medium Priority
?
579 Views
Last Modified: 2012-12-29
on this web page http://www.scie-socialcareonline.org.uk/topics.asp?guid=64f07a36-85f2-4aac-a862-61b9116190ad if we click on expand all in the list of browse topics. How can we extract all the hyperlinks of the with titles like adoption, access to birth records etc
0
Comment
Question by:mmalik15
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 4
  • 3
7 Comments
 
LVL 75

Accepted Solution

by:
käµfm³d   👽 earned 2000 total points
ID: 38729997
The subsections are displayed by simply changing the display style from none to block, and the links exist in the source HTML (i.e. they are not pulled via AJAX). For this reason you should be able to just select all the links within that section.

If you're still using Html Agility Pack, then you could do:

doc.DocumentNode.SelectNodes("//span[@class='branch']//a[not(starts-with(@href, 'javascript:'))]")

Open in new window

0
 

Author Comment

by:mmalik15
ID: 38730014
Many thanks again kaufmed..

how can i exclude rss link in the xpath? Apart from that its working fine.

Also could you kindly tell me any xpath tool to extract the information from html DOM or what's the best approach to write xpath for html dom?
0
 
LVL 75

Expert Comment

by:käµfm³d 👽
ID: 38730026
Oh, sorry. I meant to exclude that as well:

doc.DocumentNode.SelectNodes("//span[@class='branch']//a[not(starts-with(@href, 'javascript:')) and not(starts-with(@href, 'rss/'))]")

Open in new window

0
Nothing ever in the clear!

This technical paper will help you implement VMware’s VM encryption as well as implement Veeam encryption which together will achieve the nothing ever in the clear goal. If a bad guy steals VMs, backups or traffic they get nothing.

 

Author Comment

by:mmalik15
ID: 38730035
Brilliant kaufmed. Its working perfectly.

I use Altova to test any xpath on xml documents but wonder if  there is a similar tool to test Html DOM.
0
 
LVL 75

Expert Comment

by:käµfm³d 👽
ID: 38730041
I don't know of any. HTML is becoming more in line with XML with new standards that are released. Most of the frameworks people use today to build HTML do so such that the HTML is well-formed (similar to XML). As such, you should be able to use Altova on any well-formed HTML since HTML is (technically) a subset of XML (even though HTML was around first). Unless you are dealing with someone who hand-code their web page, you should be OK using Altova.
0
 
LVL 75

Expert Comment

by:käµfm³d 👽
ID: 38730044
P.S.

One of the reasons HTML Agility Pack is so popular is that the team sought to make a library that could handle (as best as one can) mal-formed HTML. HAP takes some liberties in making the source HTML well-formed so that you can use XPath against the loaded document.
0
 

Author Closing Comment

by:mmalik15
ID: 38730054
Thanks kaufmed... Its worth having EE membership because of the presence of people like you!
0

Featured Post

Veeam Task Manager for Hyper-V

Task Manager for Hyper-V provides critical information that allows you to monitor Hyper-V performance by displaying real-time views of CPU and memory at the individual VM-level, so you can quickly identify which VMs are using host resources.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Just a quick little trick I learned recently.  Now that I'm using jQuery with abandon in my asp.net applications, I have grown tired of the following syntax:      (CODE) I suppose it just offends my sense of decency to put inline VBScript on a…
ASP.Net to Oracle Connectivity Recently I had to develop an ASP.NET application connecting to an Oracle database.As I am doing it first time ,I had to solve several problems. This article will help to such developers  to develop an ASP.NET client…
Video by: ITPro.TV
In this episode Don builds upon the troubleshooting techniques by demonstrating how to properly monitor a vSphere deployment to detect problems before they occur. He begins the show using tools found within the vSphere suite as ends the show demonst…
Please read the paragraph below before following the instructions in the video — there are important caveats in the paragraph that I did not mention in the video. If your PaperPort 12 or PaperPort 14 is failing to start, or crashing, or hanging, …

656 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question