Solved

extract data from web

Posted on 2011-03-18
8
173 Views
Last Modified: 2013-11-19
Hi,

collecting data from websites manually is very hard and time consuming, I'm looking for free application to extract data from web and put it in database or file. for example: I need to get the university information (faculties, department, members, contact information .. etc )

I found different application, but it is not free and hard to learn and customized. can you please guide me to find easy, free and powerful application that do the example mentioned above in a short time.

thanks
0
Comment
Question by:nmokhayesh
  • 4
  • 3
8 Comments
 
LVL 75

Expert Comment

by:Michel Plungjan
ID: 35170400
0
 
LVL 108

Expert Comment

by:Ray Paseur
ID: 35173465
Do you want to search the contents of web sites, or do you want to copy the web sites?
0
 

Author Comment

by:nmokhayesh
ID: 35173574
I need to copy selected contents
for instance
list of all professors names and their contacts (email , tel, website, research interests)
list of all departments and contact information
list of schools and programs discription in each one

thank
 
0
 
LVL 108

Expert Comment

by:Ray Paseur
ID: 35173641
I have used httrack and it worked fairly well to make a copy of the web site onto my hard drive.  Selection of the contents was still the major issue.  Although the local web site was faster than using the internet, you would still have to manually or programmatically isolate the information you wanted to keep.

It might be possible to get a copy of Wrensoft Zoom Indexer and use that to spider the site.  Caveat: I have never tried that on a site that I did not control.

One other possibility might be to contact the site owners and ask if they can isolate this information for you.  Educational institutions are often willing to help with requests like this.
0
Do You Know the 4 Main Threat Actor Types?

Do you know the main threat actor types? Most attackers fall into one of four categories, each with their own favored tactics, techniques, and procedures.

 

Author Comment

by:nmokhayesh
ID: 35216825
OK I need to extract selected data from some university web pages to XML or excel sheet file using web scraping software

can you please tell me which free web scraping application can do this job in easy way. I search it but i got a lot of application but I do not know which one is useful/efficient

Thanks
Naif
0
 
LVL 108

Accepted Solution

by:
Ray Paseur earned 500 total points
ID: 35220096
There is no "easy way" because there is no clear vision of what you want to extract.  Each university web site is likely to be a bespoke application, so each such scraping and extraction algorithm will require custom programming.

That is why I recommended that you contact the site owners and ask if they can isolate this information for you.
0
 

Author Closing Comment

by:nmokhayesh
ID: 36710149
still not solved 100%
0
 
LVL 108

Expert Comment

by:Ray Paseur
ID: 36710198
No, it will never get solved 100%, full stop.  Here is what you asked for back in March (how many months ago was that?)

...find easy, free and powerful application that do the example mentioned above in a short time.

It would surprise me if you find easy, free and powerful all in the same package.  Those things are like Ohm's law.  Fix any two variables and the third is determined.
0

Featured Post

How your wiki can always stay up-to-date

Quip doubles as a “living” wiki and a project management tool that evolves with your organization. As you finish projects in Quip, the work remains, easily accessible to all team members, new and old.
- Increase transparency
- Onboard new hires faster
- Access from mobile/offline

Join & Write a Comment

Suggested Solutions

When setting up new project requests for our site, one of the most powerful tools our team has available to use is Axure (http://www.axure.com/). It’s a tool for creating software and web prototypes that can function and interact as if it were the a…
Password hashing is better than message digests or encryption, and you should be using it instead of message digests or encryption.  Find out why and how in this article, which supplements the original article on PHP Client Registration, Login, Logo…
Any person in technology especially those working for big companies should at least know about the basics of web accessibility. Believe it or not there are even laws in place that require businesses to provide such means for the disabled and aging p…
The viewer will get a basic understanding of what section 508 compliance can entail, learn about skip navigation links, alt text, transcripts, and font size controls.

747 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

13 Experts available now in Live!

Get 1:1 Help Now