Solved

Perl Script Modification

Posted on 2012-03-12
10
325 Views
Last Modified: 2012-03-19
Hi,

I'm currently working with two scripts.
Running pipeline.sh generates three result files in the ~/candidates folder:

'proteome'-statistics file
'proteome'-final file
'proteome'-endfile file.

It is the generated statistics file I am interested in (see attached example).
I am interested in the number of "Molecular mimicry 14-mers":
statistics file:

Parasite proteome	0
Parasite-specific proteins	0
Conserved proteins	0
Parasite 14-mers	0
Parasite-specific 14-mers	0
Molecular mimicry candidate 14-mers	23
Molecular mimicry candidate proteins	9



I want to modify the script so that none of the above files are generated.
Instead the script will create a file called "Results"
As I say, I am only interested in Molecular Mimicry 14-mers
so, instead of creating the three files, I would like a "Results" file to be created in the ~/candidates folder. And in this file would be the number 23 (or whatever the number of molecular mimicry candidates would be).

Also, If the same script is run twice, the next result should append to the previous result in the same file (i.e. 23 [next result here])
Or if the script is run 1000 times, you'd generate 1000 results.

Any advice would be much appreciated.

Thanks.
pipeline.sh
get-candidates.txt
B-burgdorferi-statistics
0
Comment
Question by:StephenMcGowan
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 4
  • 2
  • 2
10 Comments
 
LVL 29

Expert Comment

by:Jan Springer
ID: 37714694
Can you post statistics.pl?
0
 

Author Comment

by:StephenMcGowan
ID: 37715137
Hi Jesper,

Sorry, I thought I had attached it.
It should be attached to this reply.

Thanks again,
Stephen.
statistics.txt
0
 
LVL 29

Expert Comment

by:Jan Springer
ID: 37715634
I think that this might do it.  You need the "-final" file generated as input to your counters but I've added a line to remove it when it's done being used.


#!/usr/bin/perl -w
use strict;

my $counter_in = 0;
my $counter_out = 0;

open (COUNTERIN, "<$ENV{TMP}/$ENV{PARASITE}-in");
while (my $line = <COUNTERIN>) {
        if ($line =~ /^>/) {
                chomp $line;
                $counter_in++;
        }
}
close (COUNTERIN);

open (COUNTEROUT, "<$ENV{TMP}/$ENV{PARASITE}-out");
while (my $line = <COUNTEROUT>) {
        if ($line =~ /^>/) {
                chomp $line;
                $counter_out++;
        }
}
close (COUNTEROUT);

my $total_proteins = $counter_in + $counter_out;

open (STATISTICS, ">>$ENV{CANDIDATES}/Results");

my %whole_proteins;
my $mmcandidates = 0;
open (FINAL, "<$ENV{CANDIDATES}/$ENV{PARASITE}-$ENV{HOST}-final");
while (my $line = <FINAL>) {
        chomp $line;
        if ($line =~ /^(.+)-AA:\d+/) {
                $mmcandidates++;
                $whole_proteins{$1} = 1;
        }
}
close FINAL;
unlink("<$ENV{CANDIDATES}/$ENV{PARASITE}-$ENV{HOST}-final");

print STATISTICS "Molecular mimicry candidate 14-mers\t$mmcandidates\n";
close STATISTICS;
0
Industry Leaders: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

 

Author Comment

by:StephenMcGowan
ID: 37717511
Hi Jesper,

Thanks for getting back to me.
You're almost there. I've run the script and it still produces three files:

1. Results
2. a "final" file (B_burgdorferi-H_sapiens-final)
3. a "endfile" file (B_burgdorferi-H_sapiens-endfile)

(see attached picture)

I only need the Results file, the other two files shouldn't be generated (I'm trying to save data space as I'll be running this script 1000 times).

The Results file is almost right.
The file contains:

"Molecular mimicry candidate 14-mers      23"

I wanted just the number "23" to be in the file, without the "Molecular mimicry candidate 14-mers" statement beforehand. Is this possible?

What I'm trying to achieve is so, that if I run the same script 1000 times, 1000 result numbers will append to the same results file.

Hope this makes sense.

Thanks,

Stephen.
Results.jpg
0
 
LVL 84

Assisted Solution

by:ozo
ozo earned 150 total points
ID: 37723113
unlink("<$ENV{CANDIDATES}/$ENV{PARASITE}-$ENV{HOST}-final"); #should not contain "<"
0
 
LVL 84

Assisted Solution

by:ozo
ozo earned 150 total points
ID: 37723118
print STATISTICS "Molecular mimicry candidate 14-mers\t$mmcandidates\n";
#can be
print STATISTICS "$mmcandidates\n";
0
 
LVL 29

Expert Comment

by:Jan Springer
ID: 37724751
Those two files that are being left may be from another script.  

If all of the scripts are in the same directory,

     grep final *
     grep endfile *

Please post any other scripts that reference these file names.
0
 
LVL 29

Accepted Solution

by:
Jan Springer earned 350 total points
ID: 37724764
Thanks ozo for catching my cut/paste error.

What does this do for you:

#!/usr/bin/perl -w
use strict;

my $counter_in = 0;
my $counter_out = 0;

open (COUNTERIN, "<$ENV{TMP}/$ENV{PARASITE}-in");
while (my $line = <COUNTERIN>) {
        if ($line =~ /^>/) {
                chomp $line;
                $counter_in++;
        }
}
close (COUNTERIN);

open (COUNTEROUT, "<$ENV{TMP}/$ENV{PARASITE}-out");
while (my $line = <COUNTEROUT>) {
        if ($line =~ /^>/) {
                chomp $line;
                $counter_out++;
        }
}
close (COUNTEROUT);

my $total_proteins = $counter_in + $counter_out;

open (STATISTICS, ">>$ENV{CANDIDATES}/Results");

my %whole_proteins;
my $mmcandidates = 0;
open (FINAL, "<$ENV{CANDIDATES}/$ENV{PARASITE}-$ENV{HOST}-final");
while (my $line = <FINAL>) {
        chomp $line;
        if ($line =~ /^(.+)-AA:\d+/) {
                $mmcandidates++;
                $whole_proteins{$1} = 1;
        }
}
close FINAL;
unlink("$ENV{CANDIDATES}/$ENV{PARASITE}-$ENV{HOST}-final");

print STATISTICS "$mmcandidates\n";
close STATISTICS;
0

Featured Post

Technology Partners: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Computer science students often experience many of the same frustrations when going through their engineering courses. This article presents seven tips I found useful when completing a bachelors and masters degree in computing which I believe may he…
This article will inform Clients about common and important expectations from the freelancers (Experts) who are looking at your Gig.
The viewer will learn how to pass data into a function in C++. This is one step further in using functions. Instead of only printing text onto the console, the function will be able to perform calculations with argumentents given by the user.
Simple Linear Regression

729 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question