[Webinar] Learn how to a build a cloud-first strategyRegister Now

x
?
Solved

Unix diff command to find the minus of two text files

Posted on 2014-03-17
5
Medium Priority
?
1,508 Views
Last Modified: 2014-03-21
I want to print the content of file 1a that are not present in the file 1b. Both files contain one line of similar pattered text.
Example:
$ cat 1a
1234
3456
4567
$ cat 1b
1234
5566
9999
3456
$ grep -vf 1a 1b
5566
9999


$ grep -vf 1b 1a shows those present in 1a but not in 1b. And it works perfectly, But when the files are big (1000K+ records each), the above diff command hangs. Could you suggest workaround or alternate solution that might work? Thanks you.
0
Comment
Question by:toooki
5 Comments
 
LVL 68

Accepted Solution

by:
woolmilkporc earned 2000 total points
ID: 39935760
You could use "comm". The drawback here is that both files must be sorted.

sort 1b > 1b.sort

sort 1a | comm -2 -3 - 1b.sort

will show the records present in file 1a but not in 1b(.sort)

I don't think that your grep command actually "hangs". Probably it just takes quite a long time to complete.
0
 
LVL 85

Expert Comment

by:ozo
ID: 39935771
Your example seems to show a grep command, not a diff command.
If we can use other commands, this should work:
 perl -lne '@ARGV?$s{$_}++:$s{$_}||print' 1a 1b
0
 
LVL 68

Expert Comment

by:woolmilkporc
ID: 39935789
If you want to use "diff" then this one:

diff 1b 1a | grep "^>"

will show the lines of 1a which are not in 1b, preceeded by ">".

This one

diff 1a 1b |grep "^<"

will do the same, but the lines in question will be preceeded by "<".

This will remove the prefix:

diff 1a 1b |awk '/^</ {print $2}'
0
 
LVL 38

Expert Comment

by:Gerwin Jansen, EE MVE
ID: 39938069
Off topic comment deleted.

Gerwin Jansen
EE Topic Advisor
0
 

Author Comment

by:toooki
ID: 39946613
Many thanks to all.
sort 1b > 1b.sort
sort 1a | comm -2 -3 - 1b.sort

The above worked for me!
0

Featured Post

Important Lessons on Recovering from Petya

In their most recent webinar, Skyport Systems explores ways to isolate and protect critical databases to keep the core of your company safe from harm.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

WARNING:   If you follow the instructions here, you will wipe out your VTP and VLAN configurations.  Make sure you have backed up your switch!!! I recently had some issues with a few low-end Cisco routers (RV325) and I opened a case with Cisco TA…
Often times it's very very easy to extend a volume on a Linux instance in AWS, but impossible to shrink it. I wanted to contribute to the experts-exchange community a way of providing a procedure that works on an AWS instance. It can also be used on…
Monitoring a network: why having a policy is the best policy? Michael Kulchisky, MCSE, MCSA, MCP, VTSP, VSP, CCSP outlines the enormous benefits of having a policy-based approach when monitoring medium and large networks. Software utilized in this v…
In this brief tutorial Pawel from AdRem Software explains how you can quickly find out which services are running on your network, or what are the IP addresses of servers responsible for each service. Software used is freeware NetCrunch Tools (https…
Suggested Courses

864 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question