Solved

# Identifying (similar) Names and Addresses

Posted on 2014-02-10
274 Views
Hi,

I have 50,000 names and addresses from multiple sources.

The wil be at least 30% duplication.

I.e. One specific name and address may be there more than once but many NOT be 100% identical.

E.g.
John Smith, 1 High Street, London
John Smith, 1 The High Street London

Can anyone guide me to a utility which would identify names/addresses that are not quite 100% matched.

Any thoughts out there?
0
Question by:Patrick O'Dea

LVL 84

Accepted Solution

Scott McDaniel (Microsoft Access MVP - EE MVE ) earned 500 total points
That's going to be difficult to do.

You could use something like a Soundex algorithm to "rank" each one compared to the others. Essentially this would give you the greatest chance of duplicates for each entry, and you could then decide what to do with them.

Here's the wikipedia take on the Soundex stuff: http://en.wikipedia.org/wiki/Soundex

Essentially it involves replacing the characters in a string with numeric values, and then comparing the results. There are many different types of these algorithms, for various purposes. One example is this:

Consider the word "Cranston"

You keep the first letter ("C"), and then remove all other vowels, and any occurrence of letters y, h and w, so you're left with this:

Crnstn

You then assign values to the next 3 items. Using the wikipedia method, that would be:

C652

The letter "r" is = 6, the letter "n" is = 5 and the letter "s" = 2.

You'd do the same for all the strings (and you could go out further than 3 letters if you'd prefer), and store this value in a column in that table. You then sort by that column, and you can see immediately which strings are most closely related.

Allen browne has one here: http://allenbrowne.com/vba-Soundex.html. It uses a setup very much like what is described in the wikipedia link.
0

Author Closing Comment

Thanks, I will experiment.  (I heard about soundex years ago).
0

## Featured Post

### Suggested Solutions

APEX (Application Express) is used to develop a web application from Oracle. SQL Workshop is one of the tools that comes with Oracle APEX to query or modify the database objects or to make any changes to the structure.
Read about achieving the basic levels of HRIS security in the workplace.
In Microsoft Access, learn how to use Dlookup and other domain aggregate functions and one method of specifying a string value within a string. Specify the first argument, which is the expression to be returned: Specify the second argument, which …
In Microsoft Access, when working with VBA, learn some techniques for writing readable and easily maintained code.