6016

How to read a XML file with spark that contains multiple namespaces?

Question:

I'm using the spark-xml library in Azure-Databricks. But I can't get the options right to read this kind of file that contains multiple namespaces.

So I'm looking for some help to get this coded in the options, or any other approach.

Here is a stripped sample.

<msg:TrainTrackingMessage xmlns:msg="be:brail:nmbs-it:esb:msg:traintraffic" xmlns:trtf="be:brail:nmbs-it:esb:traintraffic" xmlns:gene="be:brail:nmbs-it:esb:generalelements"> <gene:Event> <gene:EventType>tracking</gene:EventType> <gene:EventMessage>TrainTracking</gene:EventMessage> <gene:EventTimeStamp>2018-09-27T14:13:15.458439</gene:EventTimeStamp> </gene:Event> <gene:Train> <gene:TrainKey> <gene:CirculationType>1</gene:CirculationType> <gene:Discriminator>0</gene:Discriminator> <gene:DepartureDate>2018-09-27</gene:DepartureDate> </gene:TrainKey> <gene:TrainNumberEBP>2E0xaZ12</gene:TrainNumberEBP> <gene:TrainDetails> <gene:TrainGroup>1</gene:TrainGroup> </gene:TrainDetails> </gene:Train> <trtf:TrainTracking> <gene:ItineraryPoint> <gene:PtcarIdentification>592</gene:PtcarIdentification> <gene:OrderNumber>150</gene:OrderNumber> <gene:ItineraryPointDetails> <gene:OperationCode>=</gene:OperationCode> <gene:CommercialStop>2</gene:CommercialStop> </gene:ItineraryPointDetails> <gene:ItineraryPointTimeInfo> <gene:ArrivalTime>14:10:47</gene:ArrivalTime> <gene:DepartureTime>14:10:54</gene:DepartureTime> </gene:ItineraryPointTimeInfo> <gene:ItineraryTechnicalInfo> <gene:EngineType>21</gene:EngineType> <gene:TractionCode>E</gene:TractionCode> <gene:TractionOperator/> </gene:ItineraryTechnicalInfo> </gene:ItineraryPoint> <trtf:GPSPosition> <trtf:GPSAltitude>51</trtf:GPSAltitude> </trtf:GPSPosition> <trtf:Libelle>E2412</trtf:Libelle> <trtf:TrackingPointInfo> <trtf:TrackingType>2</trtf:TrackingType> <trtf:TrackingOrigin>0</trtf:TrackingOrigin> </trtf:TrackingPointInfo> <trtf:TrackingTimeInfo> <trtf:Delay>1639</trtf:Delay> </trtf:TrackingTimeInfo> </trtf:TrainTracking>

Answer1:

If people would be looking for something familiar, this did the trick.

import xml.etree.ElementTree as ET xmlfiles = dbutils.fs.ls(storage_mount_name) ##Get attribute names (for now I took all leafs of the xml structure) firstfile = xmlfiles[0].path.replace('dbfs:','/dbfs') root = ET.parse(firstfile).getroot() attributes = [node.tag for node in root.iter() if len(node)==0] clean_attribute_names = [re.sub(r'\{.*\}', '', a) for a in attributes] #Create Dataframe and save it as csv df = pd.DataFrame(columns=clean_attribute_names, index=xmlfiles) for xf in xmlfiles: afile = xf.path.replace('dbfs:','/dbfs') root = ET.parse(afile).getroot() df.loc[afile] = [node.text for node in root.iter() if node.tag in attributes]

Recommend

  • accessing multiple values in MDC with siftingappender or gsiftingappender
  • logstash splits event field values and assign to @metadata field
  • Use where queries in JSONB datatype in Rails Postgres
  • Mongodb $ne date query not working as expected
  • NSDateFormatter always returning nil
  • Solr Date Regex Query
  • Hibernate override @transient
  • where to put XSLT in client webapp
  • Printing a selection of a JSON array data out to a page using JQuery
  • fetch details from xml tag
  • Solr, search only by time among datetime fields?
  • Sorting 3 dimensional array at 2nd level based on 3rd level values
  • Read response from CURL.exe program
  • Strptime with timezones and jq
  • SimpleXML PHP Parsing [duplicate]
  • Correct AuthnRequest to Azure AD/idp
  • Group by Day and Item Total, but Output Item Names as Keys
  • Reformat JSON string with php to morris charts formats
  • System of equations for Eigenvalues and Eigenvectors
  • Error on heroku meteor app
  • Timeout for a loop in XQuery MarkLogic
  • How to read a XML file with spark that contains multiple namespaces?
  • Apache crashes with error R6016 not enough space for thread data
  • Joda DateTime Invalid format
  • Docker and systemd - service stopping after 10 seconds
  • Disabling caching in Fluent Nhibernate for a specific override
  • R update values based on event
  • Disable validator via ajax
  • Finding path obstacles in a 2D image
  • Searching for node in py2neo
  • Overflow error with load_svmlight_file() from sklearn
  • Finding intersection between two collections in MongoDb
  • CakePHP 3: change order in dateWidget
  • mysql table locked after php crashes
  • ValueError: Found arrays with inconsistent numbers of samples
  • When to stop training in caffe?
  • there is no graph with tensorboard