{"id":180,"date":"2018-11-23T10:14:49","date_gmt":"2018-11-23T10:14:49","guid":{"rendered":"https:\/\/jan.schnasse.org\/blog\/?p=180"},"modified":"2018-11-23T10:14:49","modified_gmt":"2018-11-23T10:14:49","slug":"direct-accessing-xml-with-java","status":"publish","type":"post","link":"https:\/\/jan.schnasse.org\/blog\/2018\/11\/23\/direct-accessing-xml-with-java\/","title":{"rendered":"Direct Accessing XML with Java"},"content":{"rendered":"<h1>Motivation<\/h1>\n<p>Processing of huge XML files <a href=\"https:\/\/www.balisage.net\/Proceedings\/vol5\/html\/Probst01\/BalisageVol5-Probst01.html\">can become cumbersome<\/a> if your hardware is limited.<\/p>\n<blockquote><p>&#8222;Parsing a sample 20 MB XML document<sup class=\"fn-label\"><a id=\"d78572e36-ref\" class=\"footnoteref\" href=\"https:\/\/www.balisage.net\/Proceedings\/vol5\/html\/Probst01\/BalisageVol5-Probst01.html#d78572e36\">[1]<\/a><\/sup> containing Wikipedia document abstracts into a DOM tree using the Xerces library roughly consumes about 100 MB of RAM. Other document model implementations<sup class=\"fn-label\"><a id=\"d78572e40-ref\" class=\"footnoteref\" href=\"https:\/\/www.balisage.net\/Proceedings\/vol5\/html\/Probst01\/BalisageVol5-Probst01.html#d78572e40\">[2]<\/a><\/sup> such as Saxon&#8217;s TinyTree are more memory efficient; parsing the same document in Saxon consumes about 50 MB of memory. These numbers will vary with document contents, but generally the required memory scales linearly with document size, and is typically a single-digit multiple of the file size on disk.&#8220;<\/p><\/blockquote>\n<p><span style=\"font-size: 8pt; font-family: arial, helvetica, sans-serif;\">Probst, Martin. \u201cProcessing Arbitrarily Large XML using a Persistent DOM.\u201d 2010. <a href=\"https:\/\/www.balisage.net\/Proceedings\/vol5\/html\/Probst01\/BalisageVol5-Probst01.html\">https:\/\/www.balisage.net\/Proceedings\/vol5\/html\/Probst01\/BalisageVol5-Probst01.html<\/a><br \/>\n<\/span><\/p>\n<div><\/div>\n<p>A good way to deal with huge files is to split them into smaller ones. But sometimes <a href=\"https:\/\/stackoverflow.com\/questions\/43366566\/using-stax-to-create-index-for-xml-for-quick-access\">you don&#8217;t have that option<\/a>.<\/p>\n<p>Here is where <a href=\"https:\/\/en.wikipedia.org\/wiki\/Random_access\">Random Access<\/a> comes into play. While random access of binary files is well supported by standard Java tools, this is not\u00a0 true for higher-order text-based formats like XML.<\/p>\n<h1>The Plan<\/h1>\n<ol>\n<li>Find proper access points, by taking XML structure into account.<\/li>\n<li>Translate character offsets\u00a0 to byte offsets (take encoding into account)<\/li>\n<\/ol>\n<p>This sounds straightforward.<\/p>\n<h1>Existing Libraries<\/h1>\n<p>The StAX library offers streaming access to XML data without the need of loading a complete DOM model into memory. The library comes with an <a href=\"https:\/\/docs.oracle.com\/javase\/8\/docs\/api\/javax\/xml\/stream\/XMLStreamReader.html#getLocation--\">XMLStreamReader<\/a> offering a method <a href=\"https:\/\/docs.oracle.com\/javase\/8\/docs\/api\/javax\/xml\/stream\/XMLStreamReader.html#getLocation--\">getLocation()<\/a>.<a href=\"https:\/\/docs.oracle.com\/javase\/8\/docs\/api\/javax\/xml\/stream\/Location.html#getCharacterOffset--\">getCharacterOffset()<\/a> .<\/p>\n<p>But unfortunately this will only return <strong>character offsets<\/strong>. In order to access the file with standard java readers we need <strong>byte offsets<\/strong>. UTF-8 uses variable lengths for encoding characters.\u00a0 This means that we have to reread the whole file from the beginning to calculate the byte offset from character offset. This seems not acceptable.<\/p>\n<h1>Solution<\/h1>\n<p>In the following I will introduce a solution, based on\u00a0 a generated XML parser using <a href=\"http:\/\/www.antlr.org\/\" rel=\"nofollow noreferrer\">ANTLR4<\/a>.<\/p>\n<ol>\n<li>We will use the parser to walk through the XML file. While the parser is doing it&#8217;s work it will spit out byte offsets whenever a certain criteria is fulfilled (in the example we will search for XML-Elements with the name &#8218;page&#8216;).<\/li>\n<li>I will use the byte offsets to access the XML file and to read portions of XML into a Java bean using JAXB.<\/li>\n<\/ol>\n<div class=\"post-text\">\n<p>The Following works very well on a <a href=\"https:\/\/dumps.wikimedia.org\/dewiki\" rel=\"nofollow noreferrer\">~17GB Wikipedia dump<\/a><code>\/20170501\/dewiki-20170501-pages-articles-multistream.xml.bz2<\/code> . I still had to increase heap size using <code>-xX6GB<\/code> but compared to a DOM approach this looks much more acceptable.<\/p>\n<h2>1. Get XML Grammar<\/h2>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"shell\">cd \/tmp\ngit clone https:\/\/github.com\/antlr\/grammars-v4<\/pre>\n<h2>2. Generate Parser<\/h2>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"shell\">cd \/tmp\/grammars-v4\/xml\/\nmvn clean install<\/pre>\n<h2>3. Copy Generated Java files to your Project<\/h2>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"shell\">cp -r target\/generated-sources\/antlr4 \/path\/to\/your\/project\/gen<\/pre>\n<h2>4. Hook in with a Listener to collect character offsets<\/h2>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"java\">package stack43366566;\n\nimport java.util.ArrayList;\nimport java.util.List;\n\nimport org.antlr.v4.runtime.ANTLRFileStream;\nimport org.antlr.v4.runtime.CommonTokenStream;\nimport org.antlr.v4.runtime.tree.ParseTreeWalker;\n\nimport stack43366566.gen.XMLLexer;\nimport stack43366566.gen.XMLParser;\nimport stack43366566.gen.XMLParser.DocumentContext;\nimport stack43366566.gen.XMLParserBaseListener;\n\npublic class FindXmlOffset {\n\n    List&lt;Integer&gt; offsets = null;\n    String searchForElement = null;\n\n    public class MyXMLListener extends XMLParserBaseListener {\n        public void enterElement(XMLParser.ElementContext ctx) {\n            String name = ctx.Name().get(0).getText();\n            if (searchForElement.equals(name)) {\n                offsets.add(ctx.start.getStartIndex());\n            }\n        }\n    }\n\n    public List&lt;Integer&gt; createOffsets(String file, String elementName) {\n        searchForElement = elementName;\n        offsets = new ArrayList&lt;&gt;();\n        try {\n            XMLLexer lexer = new XMLLexer(new ANTLRFileStream(file));\n            CommonTokenStream tokens = new CommonTokenStream(lexer);\n            XMLParser parser = new XMLParser(tokens);\n            DocumentContext ctx = parser.document();\n            ParseTreeWalker walker = new ParseTreeWalker();\n            MyXMLListener listener = new MyXMLListener();\n            walker.walk(listener, ctx);\n            return offsets;\n        } catch (Exception e) {\n            throw new RuntimeException(e);\n        }\n    }\n\n    public static void main(String[] arg) {\n        System.out.println(\"Search for offsets.\");\n        List&lt;Integer&gt; offsets = new FindXmlOffset().createOffsets(\"\/tmp\/dewiki-20170501-pages-articles-multistream.xml\",\n                        \"page\");\n        System.out.println(\"Offsets: \" + offsets);\n    }\n\n}<\/pre>\n<h2>5. Result<\/h2>\n<p>Prints:<\/p>\n<p>Offsets: [2441, 10854, 30257, 51419 &#8230;.<\/p>\n<h2>6. Read from Offset Position<\/h2>\n<p>To test the code I&#8217;ve written class that reads in each wikipedia page to a java object<\/p>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"java\">@JacksonXmlRootElement\nclass Page {\n public Page(){};\n public String title;\n}<\/pre>\n<p>using basically this code<\/p>\n<pre class=\"EnlighterJSRAW\" data-enlighter-language=\"null\">private Page readPage(Integer offset, String filename) {\n        try (Reader in = new FileReader(filename)) {\n            in.skip(offset);\n            ObjectMapper mapper = new XmlMapper();\n             mapper.configure(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES, false);\n            Page object = mapper.readValue(in, Page.class);\n            return object;\n        } catch (Exception e) {\n            throw new RuntimeException(e);\n        }\n    }<\/pre>\n<h1>Download<\/h1>\n<p>Find complete <a href=\"https:\/\/github.com\/jschnasse\/overflow\/tree\/master\/src\/test\/java\/stack43366566\" rel=\"nofollow noreferrer\">example on github<\/a>.<\/p>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Motivation Processing of huge XML files can become cumbersome if your hardware is limited. &#8222;Parsing a sample 20 MB XML document[1] containing Wikipedia document abstracts into a DOM tree using the Xerces library roughly consumes about 100 MB of RAM. Other document model implementations[2] such as Saxon&#8217;s TinyTree are more memory efficient; parsing the same [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6,9,15],"tags":[],"class_list":["post-180","post","type-post","status-publish","format-standard","hentry","category-development","category-java","category-software"],"_links":{"self":[{"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/posts\/180","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/comments?post=180"}],"version-history":[{"count":0,"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/posts\/180\/revisions"}],"wp:attachment":[{"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/media?parent=180"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/categories?post=180"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/jan.schnasse.org\/blog\/wp-json\/wp\/v2\/tags?post=180"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}