Solved

PHP XML Dom and HTML Entities

Posted on 2006-07-11
7
3,213 Views
Last Modified: 2012-08-13
Hello,

I am trying to load html document into DOM using PHP and then convert it into some xml creating my own type of xml doc (using string concatenation). I only search for specific tags inside the document and those tags become part of the resultant xml (in CDATA).

I am facing problem with entities. I can load any document from web using this script. So those html documents may have any of these entities:

 
¡
¢
£
¤
....

My problem is that dom converts those entities into their printable value.
I tried setting
$doc->substituteEntities = false;

If in the document I output the xml encoding is utf-8 there is no problem it just prints xml well in browser or is saved to a file but only thing is it converts entities for example   to space. I want that it should not touch entities in document as I traverse the document. All the entities in a document are returned as XML_TEXT_NODE by dom.

So if I try to use php htmlentities($nodeValue) to convert them back to their entity equivalent it attaches meaningless characters to it. For example:

 
¡
¢

is the result when passed through htmlentities. See  added.

So this is my problem. I have tried for few hours but haven't found any solution to this.

Also is there any simpler solution to dump just everything which is inside an element rather than traversing each child node recursively?

Example output when a dom text node containing all html entities(converted to printable by dom don't know why) is passed through php's htmlentities function. Notice the ?s and à and  added:

I noticed that when I post things like ?s and boxes are converted to � here on the forum.

 
¡
¢
£
¤
Â¥
¦
§
¨
©
ª
«
¬
­
®
¯
°
±
²
³
´
µ
¶
·
¸
¹
º
»
¼
½
¾
¿
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
�
à
á
â
ã
ä
Ã¥
æ
ç
è
é
ê
ë
ì
í
î
ï
ð
ñ
ò
ó
ô
õ
ö
÷
ø
ù
ú
û
ü
ý
þ
ÿ

0
Comment
Question by:Sukhwinder Singh
  • 3
  • 2
7 Comments
 
LVL 29

Expert Comment

by:TeRReF
ID: 17082014
If you are using PHP5, give simplexml a try
http://php.net/simplexml

What you could try is to convert all ampersands to &
So:
 
would become
 

That should take care of your problem...
0
 

Author Comment

by:Sukhwinder Singh
ID: 17082376
But DOM gives me all the entities already converted as a XML_TEXT_NODE:

function dump_element ($el)
{

      global  $url;
      $nodeType = $el->nodeType;
      switch ($nodeType)
      {
            case XML_TEXT_NODE:
            {
                                                                  $fp = fopen("entities.txt", "a");
                  fwrite($fp, $el->nodeValue); // adds printable versions of enities
                  fclose($fp);
                  $temp = htmlentities($el->nodeValue, ENT_QUOTES );
                                                   // Adds strange character  and other before  
                  $str .= $temp;
                                                
                  break;
              }
...........
recursive function

This is what dom returns in entities.txt above for entities:

¡ ¢ £ ¤ ¥ ¦ § ¨ © ª « ¬ ­ ® ¯ ° ± ² ³ ´ µ ¶ · ¸ ¹ º » ¼ ½ ¾ ¿ À Á Â Ã Ä Å Æ Ç È É Ê Ë Ì Í Î Ï Ð Ñ Ò Ó Ô Õ Ö × Ø Ù Ú Û Ü Ý Þ ß à á â ã ä å æ ç è é ê ë ì í î ï ð ñ ò ó ô õ ö ÷ ø ù ú û ü ý þ ÿ

when the input in html document is:

  ¡ ¢ £ ¤ ¥ ¦ § ¨ © ª « ¬ ­ ® ¯ ° ± ² ³ ´ µ ¶ · ¸ ¹ º » ¼ ½ ¾ ¿ À Á Â Ã Ä Å Æ Ç È É Ê Ë Ì Í Î Ï Ð Ñ Ò Ó Ô Õ Ö × Ø Ù Ú Û Ü Ý Þ ß à á â ã ä å æ ç è é ê ë ì í î ï ð ñ ò ó ô õ ö ÷ ø ù ú û ü ý þ ÿ

And output when it goes through through php's htmlentites it produces:

  ¡ ¢ £ ¤ ¥ ¦ § ¨ © ª « ¬ ­ ® ¯ ° ± ² ³ ´ µ ¶ · ¸ ¹ º » ¼ ½ ¾ ¿ � � � � � � � � � � � � � � � � � � � � � � � � � � � � � � � � à á â ã ä å æ ç è é ê ë ì í î ï ð ñ ò ó ô õ ö ÷ ø ù ú û ü ý þ ÿ
0
 

Author Comment

by:Sukhwinder Singh
ID: 17089733
PLEASE DELETE THIS QUESTION. I have found the answer.
0
Better Security Awareness With Threat Intelligence

See how one of the leading financial services organizations uses Recorded Future as part of a holistic threat intelligence program to promote security awareness and proactively and efficiently identify threats.

 
LVL 29

Expert Comment

by:TeRReF
ID: 17089826
If you do not share the answer here, your point will not be refunded. Of course, when you share your solution, the points will be added to your credit again :)
0
 

Author Comment

by:Sukhwinder Singh
ID: 17098433
Because the question was going to be deleted I thought there was no benifit in posting the anwer here.

It was passing the third paramenter to htmlentities that was encoding, I passed 'utf' and result seemed to be ok.
0
 

Accepted Solution

by:
ee_ai_construct earned 0 total points
ID: 17306139
PAQ / Refund
ee ai construct, community support moderator
0

Featured Post

Better Security Awareness With Threat Intelligence

See how one of the leading financial services organizations uses Recorded Future as part of a holistic threat intelligence program to promote security awareness and proactively and efficiently identify threats.

Join & Write a Comment

This article will explain how to display the first page of your Microsoft Word documents (e.g. .doc, .docx, etc...) as images in a web page programatically. I have scoured the web on a way to do this unsuccessfully. The goal is to produce something …
Developers of all skill levels should learn to use current best practices when developing websites. However many developers, new and old, fall into the trap of using deprecated features because this is what so many tutorials and books tell them to u…
The viewer will learn how to count occurrences of each item in an array.
The viewer will learn how to look for a specific file type in a local or remote server directory using PHP.

707 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

15 Experts available now in Live!

Get 1:1 Help Now