SSIS ETL package to update/insert staging source to destination ?

Posted on 2010-08-12
Last Modified: 2013-11-30

Could someone please help me here, new to SSIS and ETL stuff, I have raw data being staged to a SQL table, which I then want to move to our production table and update/insert the record set based on a unique key.

The staging table has a row "SEQ" that identifies the record set, based on the SEQ I want to compare the destination table and if the SEQ exists, update the whole record or insert new otherwise.

I am using SSIS to acomplish this as per the image attached.

1. Connect OLE DB Source
2. Lookup transformation Editor, and configure error output to ignore failure.
3. Conditional Split
      - Insert Record Condition: ISNULL(Dest_SEQ)
      - Update Record Condition: (Dest_TYP != TYP) || (Dest_SALEDATE != SALEDATE) || (Dest_INVOICED != INVOICED) || (Dest_WEEKDATE != WEEKDATE) || (Dest_EWK != EWK) || (Dest_ZWK != ZWK) || (Dest_IY != IY) || (Dest_WK != WK) || (Dest_ROMO != ROMO) || (Dest_ROYR != ROYR)
4. New Reords go to OLE DB Destination
5. Updated records go to OLE DB Command as such:

UPDATE dbo.sales
TYP = ?,
EWK = ?,
ZWK = ?,
IY = ?,
WK = ?,
ROMO = ?,
ROYR = ?

The staging table has about 100 columns with data, so I am not sure how appropriate it is to define a Update Record condition for each one of those, perhaps my update record will be very expensive as far as resource when we are talking about 100,000 updates..

Is there a way I can change this so that in a conditional split (or using something else) to check if staging's SEQ exists in destination, if it does, update the whole record.

The above works, but ideally should I be creating a primary key on my SEQ table, and then doing an update/insert based on if the primary key exists or not, how could I accomplish this?

Right now, if I set the SEQ as a primary key on my "sales" destination table, and re-run my SSIS I get an error:

An OLE DB record is available.  Source: "Microsoft SQL Native Client"  Hresult: 0x80004005  Description: "Violation of PRIMARY KEY constraint 'PK_sales_1'. Cannot insert duplicate key in object 'dbo.sales'.".

But the table is empty, why can't it insert the SEQ key?

Thanks for all your help.

The guide I followed was:
Staging table layout:

USE [Sales]
CREATE TABLE [dbo].[stg_sales](
	[SEQ] [int] NOT NULL,
	[TYP] [varchar](1) NULL,
	[SALEDATE] [varchar](50) NULL,
	[INVOICED] [varchar](50) NULL,
	[WEEKDATE] [varchar](50) NULL,
	[EWK] [varchar](50) NULL,
	[ZWK] [varchar](50) NULL,
	[IY] [varchar](50) NULL,
	[WK] [varchar](50) NULL,
	[ROMO] [varchar](50) NULL,
	[ROYR] [varchar](50) NULL,
	[YRTD] [varchar](50) NULL,

Destination Table Layout:

USE [Sales]
CREATE TABLE [dbo].[sales](
	[SEQ] [int] NOT NULL,
	[TYP] [varchar](1) NULL,
	[SALEDATE] [varchar](50) NULL,
	[INVOICED] [varchar](50) NULL,
	[WEEKDATE] [varchar](50) NULL,
	[EWK] [varchar](50) NULL,
	[ZWK] [varchar](50) NULL,
	[IY] [varchar](50) NULL,
	[WK] [varchar](50) NULL,
	[ROMO] [varchar](50) NULL,
	[ROYR] [varchar](50) NULL,
	[YRTD] [varchar](50) NULL,

Open in new window

Question by:mirde

Accepted Solution

valkyrie_nc earned 500 total points
ID: 33423098
Instead of using a conditional split, you might try a Slowly Changing Dimension; once you walk through the wizard, it will create inserts or updates depending on whether your data exists already in the table or not.  Plus, you don't have to worry about creating your own queries. :)


LVL 30

Expert Comment

by:Reza Rad
ID: 33424545
I have a question:
do you want to look for SEQ in staging table, and if you find update it , if don't find insert?

if yest, why you don't use UNMATCH result of Lookup Transformation.
match result of lookup means rows which currently exists and need to be updated so redirect them to OLE DB Command.
un match result of lookup means rows which not exists currently and you should insert them, so redirect them to Destination.

does it make sense to you ?
please let me know if you mean anything else.

SCD ( SLowly Changing Dimenstion ) hasn't good performance, and this will be better to use other ways if possible. Of course the logic of question isn't very clear to me yet, but I think there are better ways than SCD.


Author Closing Comment

ID: 33504722
The SCD seems to work for me, as we will only be importing 10,000 - 30,000 new/updated records per day, this is a very "small" load for SQL Server/SCD.

Featured Post

Enabling OSINT in Activity Based Intelligence

Activity based intelligence (ABI) requires access to all available sources of data. Recorded Future allows analysts to observe structured data on the open, deep, and dark web.

Join & Write a Comment

Suggested Solutions

Title # Comments Views Activity
MS SQL export CSV & schedule It 9 44
Help on Setting up an identical test database 17 53
c# code 19 61
Convert int to military time 8 20
Slowly Changing Dimension Transformation component in data task flow is very useful for us to manage and control how data changes in SSIS.
Ever wondered why sometimes your SQL Server is slow or unresponsive with connections spiking up but by the time you go in, all is well? The following article will show you how to install and configure a SQL job that will send you email alerts includ…
Viewers will learn how the fundamental information of how to create a table.
Viewers will learn how to use the SELECT statement in SQL to return specific rows and columns, with various degrees of sorting and limits in place.

746 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

13 Experts available now in Live!

Get 1:1 Help Now