I've had two mysteries come up recently using etree.parse() (lxml
version 2.2.2). Can anyone shed some light on how to work around these?
Start with this input
======input.html=========
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN"
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
<html xmlns="http://www.w3.org/1999/xhtml">
<body>
<h2 class="speechTitle">
<p>Foo</p>
</h2>
</body>
</html>
======================
... then parse using ...
htmlTree = etree.parse(open("input.html"),
parser=etree.HTMLParser())
open("out.xml", "wb").write(etree.tostring(htmlTree))
... the resulting XML is corrupted in two ways:
======output.html=========
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN"
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd">
<html xmlns="http://www.w3.org/1999/xhtml" xmlns="http://www.w3.org/1999/xhtml">
<body>
<h2 class="speechTitle">
</h2>
<p>Foo</p>
</body>
</html>
=======================
Problem 1: The namespace declaration is duplicated (resulting in
invalid XML)
Problem 2: The <p> tag was moved outside of the <h2> tag.
Basic dom structure has been re-arranged.
Neither of these problems happen when I use the XMLParser (e.g.
'xmlTree = etree.parse(open("test2.html"),
parser=etree.XMLParser())'). However I don't know if my source html
will be valid XHTML or, more likely, just HTML. Is there a way to get
the HTMLParser to do the right thing?
Thanks for any suggestions or advice.