Solved

Condense DB

Posted on 2014-04-02
3
179 Views
Last Modified: 2014-04-02
I have a database of race participants.  In the past every entry was seen as a unique entity even though many were already in the database from a race they ran earlier.  Further, every participant is associated with several other tables (not all of which are FK-PK related).  I want to identify the duplicates, change the unique id in the related tables, then delete the participant from the participant table).

I would to do this via my classic asp portal so that I can re-use the utility in the future (although I am re-writing my code to check for existence when entering participants).

What's the best way to do this?  Here is my thought:
1) Do it one letter at a time (last name).
2) Order by gender, last name, first name.
3) Write existing participants to an array sorted as above.
4) Cycle through the array looking for matches (I can select the fields to look for matches on and I can see the list of participants).
5) When a match is found call a function that changes the participant id on the related tables.
6) Delete the duplicate entry from the participant table.

I will also write the utility to compare one-at-a-time and condense manually.  I just want a way to make a couple of passes taking care of the obvious ones.

What am I missing?  Is there an easier way?
0
Comment
Question by:Bob Schneider
  • 2
3 Comments
 
LVL 52

Accepted Solution

by:
Scott Fell,  EE MVE earned 500 total points
ID: 39972846
Finding duplicates is a lot harder than it seems.   Think about mis spellings, different spaces, one has a period in the name, the other does not, two people with the same name.

For deduping addresses for small databases, I will typically look through the data and test some ideas out.   Perhaps the first 3 letters of the  last name, the first 7 characters of the address, city and zip and concatenate to a new field as a key.  It's not perfect, but by taking just the first few characters, we eliminate a lot of spelling errors and by using multiple fields like address and zip it helps ensure we have the right person.

I would also test by using all small upper case.  You can use lower(mydata) to get that.


It does sound like you have a db design issue.  You shouldn't have to keep running these dedupes.

I would have 1 file of contacts and a transaction file for each race.  The transaction table would only have the ID, ContactID, RaceID, Time, Timestamp updated, Timestamp created.  When adding people to a race, you would choose contacts from the contact table and a race id from the scheduled races.

You can think of races as being just like an ecommerce transaction.  An invoice header and invoice detail.   The header in this case is the ID (raceID), event name, scheduled order, anything else.  Then the race transaction would contain the race id for linking, the contact id from the contact table and times.
0
 

Author Closing Comment

by:Bob Schneider
ID: 39973244
I agree with all points, including the one on db design.  I started this 12 years ago when I knew a lot less then I know now...not that I am an expert now.  :)  The insight is helpful.  I am piecing a script together that seems to be effective but...

Thanks a ton!
0
 
LVL 52

Expert Comment

by:Scott Fell, EE MVE
ID: 39973253
Looking at my own old code makes me cringe....
0

Featured Post

PRTG Network Monitor: Intuitive Network Monitoring

Network Monitoring is essential to ensure that computer systems and network devices are running. Use PRTG to monitor LANs, servers, websites, applications and devices, bandwidth, virtual environments, remote systems, IoT, and many more. PRTG is easy to set up & use.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Suggested Solutions

Title # Comments Views Activity
sql help 8 55
SSRS  - Dropdown with Null 3 24
Webservices in T-SQL 3 30
TSQL - How to declare table name 26 29
How to leverage one TLS certificate to encrypt Microsoft SQL traffic and Remote Desktop Services, versus creating multiple tickets for the same server.
Ever needed a SQL 2008 Database replicated/mirrored/log shipped on another server but you can't take the downtime inflicted by initial snapshot or disconnect while T-logs are restored or mirror applied? You can use SQL Server Initialize from Backup…
Microsoft Active Directory, the widely used IT infrastructure, is known for its high risk of credential theft. The best way to test your Active Directory’s vulnerabilities to pass-the-ticket, pass-the-hash, privilege escalation, and malware attacks …
A short tutorial showing how to set up an email signature in Outlook on the Web (previously known as OWA). For free email signatures designs, visit https://www.mail-signatures.com/articles/signature-templates/?sts=6651 If you want to manage em…

785 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question