?
Solved

I need a parsing script...or a command with awk, not sure

Posted on 2014-04-27
5
Medium Priority
?
298 Views
Last Modified: 2014-04-28
This is an XML file exported from my text. I need a script to parse it...as an example below this XML file

 <sms protocol="0" address="4053234432" date="1388706261627" type="2" subject="null" body="I'm heading home" toa="null" sc_toa="null" service_center="null" read="1" status="-1" locked="0" date_sent="null" readable_date="Jan 2, 2014 5:44:21 PM" contact_name="Mom" />
  <sms protocol="0" address="9726768860" date="1388786728946" type="1" subject="null" body="Tony_new_number - How are you?" toa="null" sc_toa="null" service_center="null" read="1" status="-1" locked="0" date_sent="null" readable_date="Jan 3, 2014 4:05:28 PM" contact_name="Tony Comp" />
  <sms protocol="0" address="9726768860" date="1388786847009" type="1" subject="null" body="Tony_new_number -  Fine.just anticipating getting back. I'm ready. Although I'll miss her." toa="null" sc_toa="null" service_center="null" read="1" status="-1" locked="0" date_sent="null" readable_date="Jan 3, 2014 4:07:27 PM" contact_name="Tony Comp" />

I would like it to read like the following

4053234432 "I'm heading home" Jan 2, 2014 5:44:21 "Mom"
9726768860 "Tony_new_number - How are you?" "Jan 3, 2014 4:05:28 PM" ""Tony Comp"
"9726768860" "Tony_new_number -  Fine.just anticipating getting back. I'm ready. Although I'll miss her." "Jan 3, 2014 4:07:27 PM" "Tony Comp"

So, I basically just need the:
The fields: address="xxx" body="xxxx" readable_date="xxxxx" and contact_name="xxx"

Thanks for any help
0
Comment
Question by:Viclyn
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 2
  • 2
5 Comments
 
LVL 62

Expert Comment

by:gheist
ID: 40027037
lex/bison seems more appropriate to dig randomly ordered fields you have.
0
 

Author Comment

by:Viclyn
ID: 40027140
I don't think it's very random. There are hundreds of entries, and they take on the format:

<sms

protocol=""
address=""
date=""
type=""
subject=""
body=""
toa=""
sc_toa=""
service_center=""
read=""
status=""
locked=""
date_sent=""
readable_date=""
contact_name=""

/>

I just want these four field:

address=""
body=""
readable_date=""
 contact_name=""
0
 
LVL 19

Accepted Solution

by:
simon3270 earned 2000 total points
ID: 40027248
sed 's/^.* address="\([0-9]*\). .* body=\("[^"]*"\) .* readable_date=\("[^"]*"\) contact_name=\("[^"]*"\) .*/\1 \2 \3 \4/' input_file

Open in new window


I've assumed you don't want the double quotes round the phone number (1 of your output lines has the quotes, the other 2 don't!)
0
 
LVL 19

Expert Comment

by:simon3270
ID: 40027255
My answer does assume a couple of things:
- that each entry is on one line, as in your example text
- that none of the fields contain double-quote characters.  What happens if they do?
0
 

Author Closing Comment

by:Viclyn
ID: 40027639
Very nice. Thanks
0

Featured Post

Ransomware: The New Cyber Threat & How to Stop It

This infographic explains ransomware, type of malware that blocks access to your files or your systems and holds them hostage until a ransom is paid. It also examines the different types of ransomware and explains what you can do to thwart this sinister online threat.  

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Background Still having to process all these year-end "csv" files received from all these sources (including Government entities), sometimes we have the need to examine the contents due to data error, etc... As a "Unix" shop, our only readily …
In the first part of this tutorial we will cover the prerequisites for installing SQL Server vNext on Linux.
Connecting to an Amazon Linux EC2 Instance from Windows Using PuTTY.
In a recent question (https://www.experts-exchange.com/questions/29004105/Run-AutoHotkey-script-directly-from-Notepad.html) here at Experts Exchange, a member asked how to run an AutoHotkey script (.AHK) directly from Notepad++ (aka NPP). This video…
Suggested Courses
Course of the Month11 days, 18 hours left to enroll

752 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question