[Last Call] Learn about multicloud storage options and how to improve your company's cloud strategy. Register Now

x
?
Solved

narrowing down a file

Posted on 2011-03-07
6
Medium Priority
?
268 Views
Last Modified: 2012-05-11
I have two files.

File1.txt is tab-separated and contains several fields.
File2.txt contains just one field.

I want to get a subset of file 1, such that field 3 of it matches exactly(1) one of the lines in File2.txt.  

The match has to be complete, not partial (so if one says foxnews.com/blah.html and the other says foxnews.com or vice versa - that would not be a match).  Otherwise, I could have just done grep –F –f File2.txt File1.txt.

How would I do that in a shell script?
0
Comment
Question by:aturetsky
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 3
  • 2
6 Comments
 
LVL 16

Accepted Solution

by:
sjklein42 earned 1200 total points
ID: 35064864
Should work.  Save as "joinIt.pl".

Without any test data, hard to test.  If it doesn't work, please post some test data.

# usage:   perl joinIt.pl infile.txt keyfile.txt

$infile = shift(@ARGV);
$keyfile = shift(@ARGV);

if ( ! open(INFILE, "<$infile") ) { die "*** can't open $infile: $!\n"; }
if ( ! open(KEYFILE, "<$keyfile") ) { die "*** can't open $keyfile: $!\n"; }

while ( <KEYFILE> )
{
	s/[\r\n]//g;
	$key{$_} = 1;
}

while ( <INFILE> )
{
	s/[\r\n]//g;
	@x = split(/\t/);
	if ( $key{$x[2]} ) { print "$_\n"; }
}

Open in new window

0
 
LVL 8

Expert Comment

by:point_pleasant
ID: 35069520
here is a shellscript that should work too, the delimiter in the cut commaned is a tab


for i in `cat file1 | cut -f3 -d'       '`
do
        grep -x $i file2
done
0
 
LVL 1

Author Comment

by:aturetsky
ID: 35072123
thanks, sjklein42 - it worked!

can I ask you - if I wanted to do the exact opposite and get only what's not in the keyfile - what would that look like?
0
Free Tool: IP Lookup

Get more info about an IP address or domain name, such as organization, abuse contacts and geolocation.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

 
LVL 1

Author Comment

by:aturetsky
ID: 35072688
actually, I think this might do the job (for that reverse task):


# usage:   perl joinIt.pl infile.txt keyfile.txt

$infile = shift(@ARGV);
$keyfile = shift(@ARGV);

if ( ! open(INFILE, "<$infile") ) { die "*** can't open $infile: $!\n"; }
if ( ! open(KEYFILE, "<$keyfile") ) { die "*** can't open $keyfile: $!\n"; }

while ( <KEYFILE> )
{
       s/[\r\n]//g;
       $key{$_} = 1;
}

while ( <INFILE> )
{
       s/[\r\n]//g;
       @x = split(/\t/);
       if ( not exists($key{$x[2]}) ) { print "$_\n"; }
}
0
 
LVL 8

Expert Comment

by:point_pleasant
ID: 35073184
the shell script to do the reverse would be as follows.  the echo statement is there to seperate each column 3 element from file1

for i in `cat file1 | cut -f3 -d'       '`
do
        echo ================== $i from file1 ======================
        grep -x -v $i file2
done
0
 
LVL 8

Assisted Solution

by:point_pleasant
point_pleasant earned 800 total points
ID: 35073548
sorry didn'd realize you wanted the whole line from file1.  Here is a shell script to do it.  if you want the reverse just add the -v option to grep


while read i
do
        col3=`echo $i | awk '{ print $3 }'`
        found=`grep -x -v $col3 file2`
        if [ "$found" != "" ]
        then
                echo $i
        fi
done < file1
0

Featured Post

Free Tool: SSL Checker

Scans your site and returns information about your SSL implementation and certificate. Helpful for debugging and validating your SSL configuration.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Background Still having to process all these year-end "csv" files received from all these sources (including Government entities), sometimes we have the need to examine the contents due to data error, etc... As a "Unix" shop, our only readily …
Active Directory replication delay is the cause to many problems.  Here is a super easy script to force Active Directory replication to all sites with by using an elevated PowerShell command prompt, and a tool to verify your changes.
Learn several ways to interact with files and get file information from the bash shell. ls lists the contents of a directory: Using the -a flag displays hidden files: Using the -l flag formats the output in a long list: The file command gives us mor…
In a recent question (https://www.experts-exchange.com/questions/29004105/Run-AutoHotkey-script-directly-from-Notepad.html) here at Experts Exchange, a member asked how to run an AutoHotkey script (.AHK) directly from Notepad++ (aka NPP). This video…

650 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question