?
Solved

Condense DB

Posted on 2014-04-02
3
Medium Priority
?
194 Views
Last Modified: 2014-04-02
I have a database of race participants.  In the past every entry was seen as a unique entity even though many were already in the database from a race they ran earlier.  Further, every participant is associated with several other tables (not all of which are FK-PK related).  I want to identify the duplicates, change the unique id in the related tables, then delete the participant from the participant table).

I would to do this via my classic asp portal so that I can re-use the utility in the future (although I am re-writing my code to check for existence when entering participants).

What's the best way to do this?  Here is my thought:
1) Do it one letter at a time (last name).
2) Order by gender, last name, first name.
3) Write existing participants to an array sorted as above.
4) Cycle through the array looking for matches (I can select the fields to look for matches on and I can see the list of participants).
5) When a match is found call a function that changes the participant id on the related tables.
6) Delete the duplicate entry from the participant table.

I will also write the utility to compare one-at-a-time and condense manually.  I just want a way to make a couple of passes taking care of the obvious ones.

What am I missing?  Is there an easier way?
0
Comment
Question by:Bob Schneider
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 2
3 Comments
 
LVL 53

Accepted Solution

by:
Scott Fell,  EE MVE earned 2000 total points
ID: 39972846
Finding duplicates is a lot harder than it seems.   Think about mis spellings, different spaces, one has a period in the name, the other does not, two people with the same name.

For deduping addresses for small databases, I will typically look through the data and test some ideas out.   Perhaps the first 3 letters of the  last name, the first 7 characters of the address, city and zip and concatenate to a new field as a key.  It's not perfect, but by taking just the first few characters, we eliminate a lot of spelling errors and by using multiple fields like address and zip it helps ensure we have the right person.

I would also test by using all small upper case.  You can use lower(mydata) to get that.


It does sound like you have a db design issue.  You shouldn't have to keep running these dedupes.

I would have 1 file of contacts and a transaction file for each race.  The transaction table would only have the ID, ContactID, RaceID, Time, Timestamp updated, Timestamp created.  When adding people to a race, you would choose contacts from the contact table and a race id from the scheduled races.

You can think of races as being just like an ecommerce transaction.  An invoice header and invoice detail.   The header in this case is the ID (raceID), event name, scheduled order, anything else.  Then the race transaction would contain the race id for linking, the contact id from the contact table and times.
0
 

Author Closing Comment

by:Bob Schneider
ID: 39973244
I agree with all points, including the one on db design.  I started this 12 years ago when I knew a lot less then I know now...not that I am an expert now.  :)  The insight is helpful.  I am piecing a script together that seems to be effective but...

Thanks a ton!
0
 
LVL 53

Expert Comment

by:Scott Fell, EE MVE
ID: 39973253
Looking at my own old code makes me cringe....
0

Featured Post

Windows Server 2016: All you need to know

Learn about Hyper-V features that increase functionality and usability of Microsoft Windows Server 2016. Also, throughout this eBook, you’ll find some basic PowerShell examples that will help you leverage the scripts in your environments!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Have you ever needed to get an ASP script to wait for a while? I have, just to let something else happen. Or in my case, to allow other stuff to happen while I was murdering my MySQL database with an update. The Original Issue This was written…
Naughty Me. While I was changing the database name from DB1 to DB_PROD1 (yep it's not real database name ^v^), I changed the database name and notified my application fellows that I did it. They turn on the application, and everything is working. A …
Sometimes it takes a new vantage point, apart from our everyday security practices, to truly see our Active Directory (AD) vulnerabilities. We get used to implementing the same techniques and checking the same areas for a breach. This pattern can re…
In this video, Percona Director of Solution Engineering Jon Tobin discusses the function and features of Percona Server for MongoDB. How Percona can help Percona can help you determine if Percona Server for MongoDB is the right solution for …

762 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question