Solved

How to identify the beginning of a section in a file with a keyword?

Posted on 2011-09-14
10
254 Views
Last Modified: 2012-05-12
Hi,
This is a follow up question for ID: 27285438.

In this script I missed one point. Actually, my text file includes many other things in it. And the beginning of the section that I am interested in starts with this type of line:


                         
           Submit file
    ===========================
and then here is the part that I parse

Open in new window


So can you please modify the code in other question ID: 27285438 to capture the beginning of it. And if it does not exist at all then we don't even need to parse the file.

How can I do that?

Thanks,


0
Comment
Question by:Tolgar
  • 5
  • 3
  • 2
10 Comments
 
LVL 9

Expert Comment

by:parparov
ID: 36539396
Is this grammar fixed?
i.e. optional whitespace, then Submit (capitalized) file (is it a keyword or any file name may be here),
then next line a bunch of equal signs between whitespaces?

0
 
LVL 38

Assisted Solution

by:wesly_chen
wesly_chen earned 100 total points
ID: 36539418
for ( @rippedParagraphs ) {
  my $submit_file;
   if (/\s*Submit\sfile/) {
       $submit_file=1;
   }
   if ($submit_file==1) {

        if(m|^#\s*Sandbox\s+location\s*\:\s*/sandbox/(.*?)/|) {
               $sandbox_location = $1;
               print $sandbox_location, "\n";
        }
        if (/^\#/ || !/\S/) {
            next unless $opt_flag || $file_flag;
        ...
        ...
        /^CR:/ && push(@CR, $_) && next;
        /^CS:/ && push(@CS, $_) && next;
        /^RR:/ && push(@RR, $_) && next;
        $opt_flag ? push(@Options, $_) : push(@Files, $_);
      }
  }
0
 
LVL 9

Expert Comment

by:parparov
ID: 36539627
I'd rather wait for Tolgar's grammar clarification.
0
Announcing the Most Valuable Experts of 2016

MVEs are more concerned with the satisfaction of those they help than with the considerable points they can earn. They are the types of people you feel privileged to call colleagues. Join us in honoring this amazing group of Experts.

 

Author Comment

by:Tolgar
ID: 36539632
@parparov: Yes this grammer is fixed. "file" is the keyword. Then optional white space before a bunch of equal signs.

@wesly_chen: I think you ignored the equal signs in the next line.


I am waiting for your reply,

Thanks,

 
0
 

Author Comment

by:Tolgar
ID: 36539652
So let me make it more clear:

[any number of white spaces]Submit file
[any number of white spaces][any number of equal signs] 

Open in new window

0
 
LVL 9

Expert Comment

by:parparov
ID: 36539665
Then the solution would be the slightly adjusted wesley_chen's code:
 my $submit_file = 0;
for ( @rippedParagraphs ) {
   if (/\s*Submit\s+file/) {
       $submit_file++;
       next;
   }
   if ($submit_file == 1) {
      if(/\s*\=+/) {
        $submit_file++;
      }
      else {
        $submit_file = 0; # two-line grammar didn't hold
      }
      next;
   }
   if ($submit_file == 2) {
        if(m|^#\s*Sandbox\s+location\s*\:\s*/sandbox/(.*?)/|) {
               $sandbox_location = $1;
               print $sandbox_location, "\n";
        }
        if (/^\#/ || !/\S/) {
            next unless $opt_flag || $file_flag;
        ...
        ...
        /^CR:/ && push(@CR, $_) && next;
        /^CS:/ && push(@CS, $_) && next;
        /^RR:/ && push(@RR, $_) && next;
        $opt_flag ? push(@Options, $_) : push(@Files, $_);
      }
  }
	 	
	

Open in new window

0
 
LVL 9

Accepted Solution

by:
parparov earned 400 total points
ID: 36539685
Minor polishing - the ifs should match from the beginning of the string, and garbage should not trail
 my $submit_file = 0;
for ( @rippedParagraphs ) {
   if (/^\s*Submit\s+file\s*$/) {
       $submit_file++;
       next;
   }
   if ($submit_file == 1) {
      if(/^\s*\=+\s*$/) {
        $submit_file++;
      }
      else {
        $submit_file = 0; # two-line grammar didn't hold
      }
      next;
   }
   if ($submit_file == 2) {
        if(m|^#\s*Sandbox\s+location\s*\:\s*/sandbox/(.*?)/|) {
               $sandbox_location = $1;
               print $sandbox_location, "\n";
        }
        if (/^\#/ || !/\S/) {
            next unless $opt_flag || $file_flag;
        ...
        ...
        /^CR:/ && push(@CR, $_) && next;
        /^CS:/ && push(@CS, $_) && next;
        /^RR:/ && push(@RR, $_) && next;
        $opt_flag ? push(@Options, $_) : push(@Files, $_);
      }
  }

Open in new window

0
 

Author Comment

by:Tolgar
ID: 36539689
oh you are very quick.

I was gonna add something else. This "Submit file" and equal sign pattern can be repeated in the file. So in this case I would like to capture all instances of it.

Is this code gonna do that?

Thanks,
0
 
LVL 38

Expert Comment

by:wesly_chen
ID: 36539691
Does  ”===========================“ always have 27 "=" characters?
0
 
LVL 9

Expert Comment

by:parparov
ID: 36539719
Change line 4 to
$submit_file = 1

Open in new window

Then you will catch all instances
0

Featured Post

Free Tool: Postgres Monitoring System

A PHP and Perl based system to collect and display usage statistics from PostgreSQL databases.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Suggested Solutions

Title # Comments Views Activity
Perl strange behaviour 5 73
transpose into pipe delemited 8 75
How to generate square thumbnail using perl 13 87
perl: Cleaning meta tags using RegEX 12 82
A year or so back I was asked to have a play with MongoDB; within half an hour I had downloaded (http://www.mongodb.org/downloads),  installed and started the daemon, and had a console window open. After an hour or two of playing at the command …
In the distant past (last year) I hacked together a little toy that would allow a couple of Manager types to query, preview, and extract data from a number of MongoDB instances, to their tool of choice: Excel (http://dilbert.com/strips/comic/2007-08…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…

828 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question