Solved

XSLT - Applying an XSLT aginst and HTML document to produce XML output

Posted on 2012-04-01
10
522 Views
Last Modified: 2012-04-14
Hi

Is there a recommened approach to applying an XSLT transformation against an HTML document ?

I would like to take an HTML input and produce some XML output.

I am currently using the SAXON XSLT processor in an java app.

Is SAXON the way to go ? Or are there any better approches I could take ?
0
Comment
Question by:Molko
  • 6
  • 2
  • 2
10 Comments
 
LVL 60

Expert Comment

by:Geert Bormans
ID: 37793049
In my opinion Saxon is about the best choice you could make.

Transforming HTML into XML usually implies that you need to add nested grouping from flat list
H1 H2 H1
to become
section{H1 {section{H2} ...} section{H1 ...}
For that XSLT2 grouping facilities come in very handy.
So indeed, I recommend using Saxon and XSLT2

Note that XSLT requires wellformed XML as its input format.
If you are processing HTML instead of XHTML as a source,
you will need to run TagSoup or HTMLTidy to parse the HTML before you can send it to XSLT
0
 

Author Comment

by:Molko
ID: 37793128
Hi

Yes, the source is HTML and not XHTML.

Could you provide me a simple example of the XSLT2 grouping ? Very much appreciated.

Thanks
0
 
LVL 47

Expert Comment

by:for_yan
ID: 37793659
Perhpas, this woul be an example of XSLT 2 grouping:
http://stackoverflow.com/questions/2177927/grouping-several-groups-in-xslt-2
0
 
LVL 60

Expert Comment

by:Geert Bormans
ID: 37793786
well, that is a good example (and simple to grasp) for one level grouping.
it is getting really complex and harder to swallow if you want to do this up to six levels.
(I once did one in XSLT1, and it required multiple steps, never managed to get it working properly for other than straightforward examples in a single step,
so allthough complex, my statement for XSLT2 holds)
I have a stylesheet I always use, but have not done that myself, so would like to pass the reference to you, not the stylesheet to give the author credit
(you could google for it "nesting html with XSLT2 grouping" or something like that for the xsl biglist(mullberry tech)
I will try to find the reference later tonight
0
 
LVL 47

Expert Comment

by:for_yan
ID: 37793849
Perhaps you already did it, but anyway
I combined the code from here:
http://blog.msbbc.co.uk/2007/06/simple-saxon-java-example.html
downloaded saxonb9-1-0-8j.zip from here
http://sourceforge.net/projects/saxon/files/Saxon-B/9.1.0.8/saxonb9-1-0-8j.zip/download
 and expanded it and placed saxon9.jar on the classpath


and used input files from the above link.
http://stackoverflow.com/questions/2177927/grouping-several-groups-in-xslt-2

And it worked exactly as stated there

This is the code:
import javax.xml.transform.Transformer;
import javax.xml.transform.TransformerConfigurationException;
import javax.xml.transform.TransformerException;
import javax.xml.transform.TransformerFactory;
import javax.xml.transform.stream.StreamResult;
import javax.xml.transform.stream.StreamSource;
import java.io.File;

public class SimpleSaxon {


    public static void myTransformer (String sourceID, String xslID)
throws TransformerException, TransformerConfigurationException {

        // Create a transform factory instance.
        TransformerFactory tfactory = TransformerFactory.newInstance();

 // Create a transformer for the stylesheet.
 Transformer transformer = tfactory.newTransformer(new StreamSource(new File(xslID)));

 // Transform the source XML to System.out.
 transformer.transform(new StreamSource(new File(sourceID)),
    new StreamResult(System.out));
}

    public static void main(String args[]) {

 // set the TransformFactory to use the Saxon TransformerFactoryImpl method
 System.setProperty("javax.xml.transform.TransformerFactory",
 "net.sf.saxon.TransformerFactoryImpl");


 String foo_xml = "input.xml"; //input xml
 String foo_xsl = "input.xsl"; //input xsl

 try {
         myTransformer (foo_xml, foo_xsl);
 } catch (Exception ex) {
      handleException(ex);
 }

}

  private static void handleException(Exception ex) {

     System.out.println("EXCEPTION: " + ex);
     ex.printStackTrace();
}


}

Open in new window


input.xml


<article>
  <h1>A section title here</h1>
  <p>A paragraph.</p>
  <p>Another paragraph.</p>
  <bl>Bulleted list item.</bl>
  <bl>Another bulleted list item.</bl>
  <h1>Another section title</h1>
  <p>Yet another paragraph.</p>
</article>

Open in new window


input.xsl:

<xsl:stylesheet
  version="2.0"
  xmlns:xsl="http://www.w3.org/1999/XSL/Transform">

  <xsl:strip-space elements="*"/>
  <xsl:output indent="yes"/>

  <xsl:template match="article">
    <xsl:copy>
      <xsl:for-each-group select="*" group-starting-with="h1">
        <sec>
          <xsl:copy-of select="."/>
          <xsl:for-each-group select="current-group() except ." group-adjacent="boolean(self::bl)">
            <xsl:choose>
              <xsl:when test="current-grouping-key()">
                <list>
                  <xsl:apply-templates select="current-group()"/>
                </list>
              </xsl:when>
              <xsl:otherwise>
                <xsl:copy-of select="current-group()"/>
              </xsl:otherwise>
            </xsl:choose>
          </xsl:for-each-group>
        </sec>
      </xsl:for-each-group>
    </xsl:copy>
  </xsl:template>

  <xsl:template match="bl">
    <list-item>
      <xsl:apply-templates/>
    </list-item>
  </xsl:template>

</xsl:stylesheet>

Open in new window



Output:

<?xml version="1.0" encoding="UTF-8"?>
<article>
   <sec>
      <h1>A section title here</h1>
      <p>A paragraph.</p>
      <p>Another paragraph.</p>
      <list>
         <list-item>Bulleted list item.</list-item>
         <list-item>Another bulleted list item.</list-item>
      </list>
   </sec>
   <sec>
      <h1>Another section title</h1>
      <p>Yet another paragraph.</p>
   </sec>
</article>

Open in new window

0
DevOps Toolchain Recommendations

Read this Gartner Research Note and discover how your IT organization can automate and optimize DevOps processes using a toolchain architecture.

 
LVL 60

Expert Comment

by:Geert Bormans
ID: 37793857
my friend, what I meant was more than one level.
It is easy with one level,
h1 - h1 - h1
and it is easy with predictable levels (see michael kays excellent reference book)
h1 - h2 - h2 - h1 - h 2 - h3 - h2 - h3 - h1

it is somewhat harder with unpredictable levels (such as most html out there)
h1 - h4 - h2 - h1 - h4 - h3 - h2 - h1- h4 - h3

please read my comments more carefully :-)
0
 
LVL 60

Accepted Solution

by:
Geert Bormans earned 500 total points
ID: 37798378
have a look here if you are using XSLT2
http://www.dpawson.co.uk/xsl/rev2/html.html
David Carlisle developed a XSLT2 html cleanup you can use as a first step
0
 
LVL 60

Expert Comment

by:Geert Bormans
ID: 37799574
And here is the code I use for nesting flat structures

http://stackoverflow.com/questions/2108348/xslt-deepening-content-structure

the answer from martin honnen is what you are looking for
... somewhat advanced stuff, so you will need to take some time to swallow and adapt it

cheers

Geert
0
 

Author Closing Comment

by:Molko
ID: 37846083
Thanks
0
 
LVL 60

Expert Comment

by:Geert Bormans
ID: 37846201
welcome
0

Featured Post

Is Your Active Directory as Secure as You Think?

More than 75% of all records are compromised because of the loss or theft of a privileged credential. Experts have been exploring Active Directory infrastructure to identify key threats and establish best practices for keeping data safe. Attend this month’s webinar to learn more.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Suggested Solutions

This was posted to the Netbeans forum a Feb, 2010 and I also sent it to Verisign. Who didn't help much in my struggles to get my application signed. ------------------------- Start The idea here is to target your cell phones with the correct…
Introduction This article is the second of three articles that explain why and how the Experts Exchange QA Team does test automation for our web site. This article covers the basic installation and configuration of the test automation tools used by…
This tutorial explains how to use the VisualVM tool for the Java platform application. This video goes into detail on the Threads, Sampler, and Profiler tabs.
Viewers will learn how to properly install Eclipse with the necessary JDK, and will take a look at an introductory Java program. Download Eclipse installation zip file: Extract files from zip file: Download and install JDK 8: Open Eclipse and …

910 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

21 Experts available now in Live!

Get 1:1 Help Now