• Status: Solved
  • Priority: Medium
  • Security: Public
  • Views: 392
  • Last Modified:

An Intelligent Script

Hi Experts!

I'd like create a PHP script that extract relevant information from WebPage and i need your help.

The script captures the html code od various similar page from a site (for example all articles page). I define which parts of body (i.e. title of articles) may be extract and i'd like create an algorithm that define, automatically, regular expression for that parts.

I need your suggestions.
0
ttiero
Asked:
ttiero
  • 2
1 Solution
 
aolXFTCommented:
Unless you show us a sample of the body you want to extract information from we can't really help.

From your question though, I'd consider using XML functions. You might have to run your document through tidy first though to make it XHTML(and therefore XML) Compliant. You can install tidy as a PHP/PECL Extension ( http://pecl.php.net/package/tidy ).

Then use XML Functions to get what you want.
0
 
ttieroAuthor Commented:
There is PHP classes that transform HTML to XML?
0
 
aolXFTCommented:
There is an PHP Extension for Tidy. you can check out tidy at http://www.w3.org/People/Raggett/tidy/.  John Coggeshall wrote a PHP Extension for libTidy, or tidyLib, or whatever it is, and submitted it to PECL, check out the above url, or http://www.coggeshall.org/tidy.php

The command line tidy client allows the switch -asxml so I'm sure you can do the same with the PHP extension.

That may however be overkill depending on how complex your document is. Basicly I need a sample.
0

Featured Post

Free Tool: IP Lookup

Get more info about an IP address or domain name, such as organization, abuse contacts and geolocation.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

  • 2
Tackle projects and never again get stuck behind a technical roadblock.
Join Now