Solved

Condense DB

Posted on 2014-04-02
3
175 Views
Last Modified: 2014-04-02
I have a database of race participants.  In the past every entry was seen as a unique entity even though many were already in the database from a race they ran earlier.  Further, every participant is associated with several other tables (not all of which are FK-PK related).  I want to identify the duplicates, change the unique id in the related tables, then delete the participant from the participant table).

I would to do this via my classic asp portal so that I can re-use the utility in the future (although I am re-writing my code to check for existence when entering participants).

What's the best way to do this?  Here is my thought:
1) Do it one letter at a time (last name).
2) Order by gender, last name, first name.
3) Write existing participants to an array sorted as above.
4) Cycle through the array looking for matches (I can select the fields to look for matches on and I can see the list of participants).
5) When a match is found call a function that changes the participant id on the related tables.
6) Delete the duplicate entry from the participant table.

I will also write the utility to compare one-at-a-time and condense manually.  I just want a way to make a couple of passes taking care of the obvious ones.

What am I missing?  Is there an easier way?
0
Comment
Question by:Bob Schneider
  • 2
3 Comments
 
LVL 52

Accepted Solution

by:
Scott Fell,  EE MVE earned 500 total points
ID: 39972846
Finding duplicates is a lot harder than it seems.   Think about mis spellings, different spaces, one has a period in the name, the other does not, two people with the same name.

For deduping addresses for small databases, I will typically look through the data and test some ideas out.   Perhaps the first 3 letters of the  last name, the first 7 characters of the address, city and zip and concatenate to a new field as a key.  It's not perfect, but by taking just the first few characters, we eliminate a lot of spelling errors and by using multiple fields like address and zip it helps ensure we have the right person.

I would also test by using all small upper case.  You can use lower(mydata) to get that.


It does sound like you have a db design issue.  You shouldn't have to keep running these dedupes.

I would have 1 file of contacts and a transaction file for each race.  The transaction table would only have the ID, ContactID, RaceID, Time, Timestamp updated, Timestamp created.  When adding people to a race, you would choose contacts from the contact table and a race id from the scheduled races.

You can think of races as being just like an ecommerce transaction.  An invoice header and invoice detail.   The header in this case is the ID (raceID), event name, scheduled order, anything else.  Then the race transaction would contain the race id for linking, the contact id from the contact table and times.
0
 

Author Closing Comment

by:Bob Schneider
ID: 39973244
I agree with all points, including the one on db design.  I started this 12 years ago when I knew a lot less then I know now...not that I am an expert now.  :)  The insight is helpful.  I am piecing a script together that seems to be effective but...

Thanks a ton!
0
 
LVL 52

Expert Comment

by:Scott Fell, EE MVE
ID: 39973253
Looking at my own old code makes me cringe....
0

Featured Post

Better Security Awareness With Threat Intelligence

See how one of the leading financial services organizations uses Recorded Future as part of a holistic threat intelligence program to promote security awareness and proactively and efficiently identify threats.

Join & Write a Comment

Use this article to create a batch file to backup a Microsoft SQL Server database to a Windows folder.  The folder can be on the local hard drive or on a network share.  This batch file will query the SQL server to get the current date & time and wi…
In this article we will get to know that how can we recover deleted data if it happens accidently. We really can recover deleted rows if we know the time when data is deleted by using the transaction log.
Sending a Secure fax is easy with eFax Corporate (http://www.enterprise.efax.com). First, Just open a new email message.  In the To field, type your recipient's fax number @efaxsend.com. You can even send a secure international fax — just include t…
Internet Business Fax to Email Made Easy - With eFax Corporate (http://www.enterprise.efax.com), you'll receive a dedicated online fax number, which is used the same way as a typical analog fax number. You'll receive secure faxes in your email, fr…

708 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

17 Experts available now in Live!

Get 1:1 Help Now