SSIS ETL package to update/insert staging source to destination ?

Posted on 2010-08-12
Last Modified: 2013-11-30

Could someone please help me here, new to SSIS and ETL stuff, I have raw data being staged to a SQL table, which I then want to move to our production table and update/insert the record set based on a unique key.

The staging table has a row "SEQ" that identifies the record set, based on the SEQ I want to compare the destination table and if the SEQ exists, update the whole record or insert new otherwise.

I am using SSIS to acomplish this as per the image attached.

1. Connect OLE DB Source
2. Lookup transformation Editor, and configure error output to ignore failure.
3. Conditional Split
      - Insert Record Condition: ISNULL(Dest_SEQ)
      - Update Record Condition: (Dest_TYP != TYP) || (Dest_SALEDATE != SALEDATE) || (Dest_INVOICED != INVOICED) || (Dest_WEEKDATE != WEEKDATE) || (Dest_EWK != EWK) || (Dest_ZWK != ZWK) || (Dest_IY != IY) || (Dest_WK != WK) || (Dest_ROMO != ROMO) || (Dest_ROYR != ROYR)
4. New Reords go to OLE DB Destination
5. Updated records go to OLE DB Command as such:

UPDATE dbo.sales
TYP = ?,
EWK = ?,
ZWK = ?,
IY = ?,
WK = ?,
ROMO = ?,
ROYR = ?

The staging table has about 100 columns with data, so I am not sure how appropriate it is to define a Update Record condition for each one of those, perhaps my update record will be very expensive as far as resource when we are talking about 100,000 updates..

Is there a way I can change this so that in a conditional split (or using something else) to check if staging's SEQ exists in destination, if it does, update the whole record.

The above works, but ideally should I be creating a primary key on my SEQ table, and then doing an update/insert based on if the primary key exists or not, how could I accomplish this?

Right now, if I set the SEQ as a primary key on my "sales" destination table, and re-run my SSIS I get an error:

An OLE DB record is available.  Source: "Microsoft SQL Native Client"  Hresult: 0x80004005  Description: "Violation of PRIMARY KEY constraint 'PK_sales_1'. Cannot insert duplicate key in object 'dbo.sales'.".

But the table is empty, why can't it insert the SEQ key?

Thanks for all your help.

The guide I followed was:
Staging table layout:

USE [Sales]
CREATE TABLE [dbo].[stg_sales](
	[SEQ] [int] NOT NULL,
	[TYP] [varchar](1) NULL,
	[SALEDATE] [varchar](50) NULL,
	[INVOICED] [varchar](50) NULL,
	[WEEKDATE] [varchar](50) NULL,
	[EWK] [varchar](50) NULL,
	[ZWK] [varchar](50) NULL,
	[IY] [varchar](50) NULL,
	[WK] [varchar](50) NULL,
	[ROMO] [varchar](50) NULL,
	[ROYR] [varchar](50) NULL,
	[YRTD] [varchar](50) NULL,

Destination Table Layout:

USE [Sales]
CREATE TABLE [dbo].[sales](
	[SEQ] [int] NOT NULL,
	[TYP] [varchar](1) NULL,
	[SALEDATE] [varchar](50) NULL,
	[INVOICED] [varchar](50) NULL,
	[WEEKDATE] [varchar](50) NULL,
	[EWK] [varchar](50) NULL,
	[ZWK] [varchar](50) NULL,
	[IY] [varchar](50) NULL,
	[WK] [varchar](50) NULL,
	[ROMO] [varchar](50) NULL,
	[ROYR] [varchar](50) NULL,
	[YRTD] [varchar](50) NULL,

Open in new window

Question by:mirde
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions

Accepted Solution

valkyrie_nc earned 500 total points
ID: 33423098
Instead of using a conditional split, you might try a Slowly Changing Dimension; once you walk through the wizard, it will create inserts or updates depending on whether your data exists already in the table or not.  Plus, you don't have to worry about creating your own queries. :)


LVL 30

Expert Comment

by:Reza Rad
ID: 33424545
I have a question:
do you want to look for SEQ in staging table, and if you find update it , if don't find insert?

if yest, why you don't use UNMATCH result of Lookup Transformation.
match result of lookup means rows which currently exists and need to be updated so redirect them to OLE DB Command.
un match result of lookup means rows which not exists currently and you should insert them, so redirect them to Destination.

does it make sense to you ?
please let me know if you mean anything else.

SCD ( SLowly Changing Dimenstion ) hasn't good performance, and this will be better to use other ways if possible. Of course the logic of question isn't very clear to me yet, but I think there are better ways than SCD.


Author Closing Comment

ID: 33504722
The SCD seems to work for me, as we will only be importing 10,000 - 30,000 new/updated records per day, this is a very "small" load for SQL Server/SCD.

Featured Post

NEW Veeam Agent for Microsoft Windows

Backup and recover physical and cloud-based servers and workstations, as well as endpoint devices that belong to remote users. Avoid downtime and data loss quickly and easily for Windows-based physical or public cloud-based workloads!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

International Data Corporation (IDC) prognosticates that before the current the year gets over disbursing on IT framework products to be sent in cloud environs will be $37.1B.
It is possible to export the data of a SQL Table in SSMS and generate INSERT statements. It's neatly tucked away in the generate scripts option of a database.
Using examples as well as descriptions, and references to Books Online, show the different Recovery Models available in SQL Server and explain, as well as show how full, differential and transaction log backups are performed
This videos aims to give the viewer a basic demonstration of how a user can query current session information by using the SYS_CONTEXT function

626 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question