Want to win a PS4? Go Premium and enter to win our High-Tech Treats giveaway. Enter to Win

x
?
Solved

URL stream to xhtml

Posted on 2009-07-16
10
Medium Priority
?
378 Views
Last Modified: 2012-05-07
How would I modify the below to convert a URL stream into an xhtml string, rather than an html file into an xhtml file?
import org.w3c.tidy.Tidy;
import java.io.FileInputStream;
import java.io.FileOutputStream;
import org.w3c.dom.Document;
 
public class HTML_to_XHTML{
   public static void main(String[] args){
      try{
         FileInputStream FIS=new FileInputStream("C://test.html");
         FileOutputStream FOS=new FileOutputStream("C://testXHTML.xml");   
         Tidy T=new Tidy();
         Document D=T.parseDOM(FIS,FOS);
         }
      catch (java.io.FileNotFoundException e)
         {System.out.println(e.getMessage());}   
      }
   }
}

Open in new window

0
Comment
Question by:arichexe
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 5
  • 4
10 Comments
 
LVL 86

Expert Comment

by:CEHJ
ID: 24871675
0
 
LVL 92

Expert Comment

by:objects
ID: 24875255
        InputStream FIS=url.getInputStream();
         StringWriter FOS=new StringWriter();  
         Tidy T=new Tidy();
         T.parseDOM(FIS,FOS);
         String xhtml = FOS.toString();
0
 

Author Comment

by:arichexe
ID: 24899921
I'm getting a "Tidy cannot be resolved to a type" error.
<%@ page import="java.io.*,java.net.*,java.text.*,java.util.*,javax.xml.parsers.*,javax.xml.xpath.*,org.w3c.dom.*,org.w3c.dom.*,org.xml.sax.*,org.w3c.tidy.*" %>
<%
URL url = new URL(MyUrl);
HttpURLConnection conn = (HttpURLConnection) url.openConnection();
conn.setRequestMethod("POST");
conn.setRequestProperty("Content-Type","text/xml");
conn.setDoOutput(true);
OutputStream os = conn.getOutputStream();
os.flush();
os.close();
 
InputStream is = conn.getInputStream();
StringWriter ox = new StringWriter();
Tidy T=new Tidy();
T.parseDOM(is,ox);
String xhtml = ox.toString();
 
DocumentBuilderFactory factory = DocumentBuilderFactory.newInstance();
factory.setValidating(false);
factory.setIgnoringElementContentWhitespace(true);
DocumentBuilder builder = factory.newDocumentBuilder();
Document document = builder.parse(new InputSource(new StringReader(xhtml)));
document.getDocumentElement().normalize();
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodeList = (NodeList) xpath.evaluate("//title/text()",document,XPathConstants.NODESET);
 
if (nodeList.getLength() > 0) {
  for (int i = 0; i < nodeList.getLength(); i++) {
    out.print("msg: " + nodeList.item(i).toString());
  }
}else{
  out.print("msg: not found");
}
%>

Open in new window

0
What does it mean to be "Always On"?

Is your cloud always on? With an Always On cloud you won't have to worry about downtime for maintenance or software application code updates, ensuring that your bottom line isn't affected.

 
LVL 92

Expert Comment

by:objects
ID: 24900113
make sure you have the tidy jar in your webapps lib directory
0
 
LVL 92

Expert Comment

by:objects
ID: 24900125
you should also be able to simplify your code to the following

InputStream is = conn.getInputStream();
StringWriter ox = new StringWriter();
Tidy T=new Tidy();
Document document = T.parseDOM(is,ox);
String xhtml = ox.toString();
document.getDocumentElement().normalize();
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodeList = (NodeList) xpath.evaluate("//title/text()",document,XPathConstants.NODESET);
 
if (nodeList.getLength() > 0) {
  for (int i = 0; i < nodeList.getLength(); i++) {
    out.print("msg: " + nodeList.item(i).toString());
  }
}else{
  out.print("msg: not found");
}
0
 

Author Comment

by:arichexe
ID: 24900831
Now I'm getting "The method parseDOM(InputStream, OutputStream) in the type Tidy is not applicable for the arguments (InputStream, StringWriter)."
<%@ page import="java.io.*,java.net.*,java.text.*,java.util.*,javax.xml.parsers.*,javax.xml.xpath.*,org.w3c.dom.*,org.w3c.dom.*,org.w3c.tidy.*,org.xml.sax.*" %>
<%
URL url = new URL(MyUrl);
HttpURLConnection conn = (HttpURLConnection) url.openConnection();
conn.setRequestMethod("POST");
conn.setRequestProperty("Content-Type","text/html");
conn.setDoOutput(true);
OutputStream os = conn.getOutputStream();
os.flush();
os.close();
 
InputStream is = conn.getInputStream();
StringWriter ox = new StringWriter();
Tidy T=new Tidy();
Document document = T.parseDOM(is,ox);
String xhtml = ox.toString();
document.getDocumentElement().normalize();
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodeList = (NodeList) xpath.evaluate("//title/text()",document,XPathConstants.NODESET);
 
if (nodeList.getLength() > 0) {
  for (int i = 0; i < nodeList.getLength(); i++) {
    out.print("msg: " + nodeList.item(i).toString());
  }
}else{
  out.print("msg: not found");
}
%>

Open in new window

0
 
LVL 92

Expert Comment

by:objects
ID: 24901038
you don't actually need to create the string at all, try this:

InputStream is = conn.getInputStream();
Tidy T=new Tidy();
Document document = T.parseDOM(is, null);
document.getDocumentElement().normalize();
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodeList = (NodeList) xpath.evaluate("//title/text()",document,XPathConstants.NODESET);
0
 

Author Comment

by:arichexe
ID: 24901110
Now it returns this weird string "msg: org.w3c.tidy.DOMTextImpl@9674b2d" and the last 7 chars change when I hit refresh.  No error, though.  Strange.
<%@ page import="java.io.*,java.net.*,java.text.*,java.util.*,javax.xml.parsers.*,javax.xml.xpath.*,org.w3c.dom.*,org.w3c.dom.*,org.w3c.tidy.*,org.xml.sax.*" %>
<%
URL url = new URL("http://MyUrl.com");
HttpURLConnection conn = (HttpURLConnection) url.openConnection();
InputStream is = conn.getInputStream();
Tidy T=new Tidy();
Document document = T.parseDOM(is,null);
document.getDocumentElement().normalize();
XPath xpath = XPathFactory.newInstance().newXPath();
NodeList nodeList = (NodeList) xpath.evaluate("//title/text()",document,XPathConstants.NODESET);
 
if (nodeList.getLength() > 0) {
  for (int i = 0; i < nodeList.getLength(); i++) {
    out.print("msg: " + nodeList.item(i).toString());
  }
}else{
  out.print("msg: not found");
}
%>

Open in new window

0
 
LVL 92

Accepted Solution

by:
objects earned 2000 total points
ID: 24901121
>     out.print("msg: " + nodeList.item(i).toString());

thats because Node doesn't have a toString(), try instead getNodeValue()

    out.print("msg: " + nodeList.item(i).getNodeValue());
0
 

Author Closing Comment

by:arichexe
ID: 31604332
Thanks!
0

Featured Post

Free Tool: Path Explorer

An intuitive utility to help find the CSS path to UI elements on a webpage. These paths are used frequently in a variety of front-end development and QA automation tasks.

One of a set of tools we're offering as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

An old method to applying the Singleton pattern in your Java code is to check if a static instance, defined in the same class that needs to be instantiated once and only once, is null and then create a new instance; otherwise, the pre-existing insta…
By the end of 1980s, object oriented programming using languages like C++, Simula69 and ObjectPascal gained momentum. It looked like programmers finally found the perfect language. C++ successfully combined the object oriented principles of Simula w…
Viewers will learn about arithmetic and Boolean expressions in Java and the logical operators used to create Boolean expressions. We will cover the symbols used for arithmetic expressions and define each logical operator and how to use them in Boole…
Viewers will learn one way to get user input in Java. Introduce the Scanner object: Declare the variable that stores the user input: An example prompting the user for input: Methods you need to invoke in order to properly get  user input:
Suggested Courses

609 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question