An Intelligent Script

Hi Experts!

I'd like create a PHP script that extract relevant information from WebPage and i need your help.

The script captures the html code od various similar page from a site (for example all articles page). I define which parts of body (i.e. title of articles) may be extract and i'd like create an algorithm that define, automatically, regular expression for that parts.

I need your suggestions.
Who is Participating?
aolXFTConnect With a Mentor Commented:
There is an PHP Extension for Tidy. you can check out tidy at  John Coggeshall wrote a PHP Extension for libTidy, or tidyLib, or whatever it is, and submitted it to PECL, check out the above url, or

The command line tidy client allows the switch -asxml so I'm sure you can do the same with the PHP extension.

That may however be overkill depending on how complex your document is. Basicly I need a sample.
Unless you show us a sample of the body you want to extract information from we can't really help.

From your question though, I'd consider using XML functions. You might have to run your document through tidy first though to make it XHTML(and therefore XML) Compliant. You can install tidy as a PHP/PECL Extension ( ).

Then use XML Functions to get what you want.
ttieroAuthor Commented:
There is PHP classes that transform HTML to XML?
Question has a verified solution.

Are you are experiencing a similar issue? Get a personalized answer when you ask a related question.

Have a better answer? Share it in a comment.

All Courses

From novice to tech pro — start learning today.