Solved

Identify outliers through SQL

Posted on 2015-02-24
10
23 Views
Last Modified: 2016-06-19
Hi- I have a large recordset of about 300,000 records and am looking to capture the outliers of a column within SQL.  

For the sake of simplicity, let's call my table mainTable and the data in question pData.

I understand that I'll need to capture the quartiles, but am not quite sure how to go about doing this.

Thanks in advance.
0
Comment
Question by:Andrew Luedke
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
10 Comments
 
LVL 44

Expert Comment

by:AndyAinscow
ID: 40628453
You could display the max and min values.  That will give you the range and then do a select query with >= or <= against the values you want.
SELECT MAX(x) as MaxX, MIN(x) as MinX FROM tbl1
then eg.
SELECT * FROM tbl1 WHERE (x >= (MaxX-5))
and
SELECT * FROM tbl1 WHERE (x <= (MinX+5))
0
 
LVL 44

Expert Comment

by:AndyAinscow
ID: 40628458
ps.  Exactly what you define as an 'outlier' is something you will have to decide.
0
 

Author Comment

by:Andrew Luedke
ID: 40628473
In terms of outliers, I'd like to grab the 1% and 99% percentiles to capture the extreme values.  Does this help to refine the algorithm?
0
The Eight Noble Truths of Backup and Recovery

How can IT departments tackle the challenges of a Big Data world? This white paper provides a roadmap to success and helps companies ensure that all their data is safe and secure, no matter if it resides on-premise with physical or virtual machines or in the cloud.

 
LVL 44

Expert Comment

by:AndyAinscow
ID: 40628500
Do you mean the 300 highest and 300 lowest or those higher than (lowest + 0.99*range)
0
 

Author Comment

by:Andrew Luedke
ID: 40628550
Apologies for the lack of clarity here.  The numbers could vary.

The formula should determine the general distribution.  Meaning that if you have a range of numbers, the formula should determine the thresholds and capture 1% of values below the normal range and the other 1% of values above.  This way, you can dynamically captures the extreme highs and lows of a set.  

For example, let's say we have a 200,000 numbers with the following characteristics:

Min:
-100

Max:
1.09

Avg:
.2

STDev:
.5

How do we go about capturing those outliers within the set?
0
 
LVL 44

Expert Comment

by:AndyAinscow
ID: 40628725
Do you want everything done in SQL or can you perform some calculations outside of the SQL to determine the limits which you then feed back into an SQL select query?
0
 

Author Comment

by:Andrew Luedke
ID: 40628735
Given the size of the database, it would be best to try and accomplish everything via SQL.
0
 
LVL 48

Accepted Solution

by:
PortletPaul earned 500 total points
ID: 40629613
without so much as a table name or field name to go by all I can do is suggest NTILE()
e.g.

SELECT
   *
  , NTILE(10) OVER (PARTITION BY [Subject] ORDER BY Marks DESC) AS [TileNo]
FROM Students

You could perhaps also "do this in both directions" so use NTILE() twice , but order ASC in one and DESC in the other, then you can filter out the outliers that have 1 one either of those columns. Also note you can alter the number of "tiles" in my example I used 10

for ranking functions in SQL 2008 see: https://msdn.microsoft.com/en-us/library/ms189798(v=sql.100).aspx

---
if you had (or have) SQL 2012 you could use PERCENT_RANK()
0
 
LVL 49

Expert Comment

by:Vitor Montalvão
ID: 40639280
Andrew, you still have the issue or it's already solved?
0

Featured Post

Free Tool: SSL Checker

Scans your site and returns information about your SSL implementation and certificate. Helpful for debugging and validating your SSL configuration.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Suggested Solutions

Title # Comments Views Activity
SQL Instance service gone? 5 39
Applying Roles in Common Scenarios 3 22
MS SQL Server Management Studio R2 4 33
Need to trim my database size 9 29
A long time ago (May 2011), I have written an article showing you how to create a DLL using Visual Studio 2005 to be hosted in SQL Server 2005. That was valid at that time and it is still valid if you are still using these versions. You can still re…
How to leverage one TLS certificate to encrypt Microsoft SQL traffic and Remote Desktop Services, versus creating multiple tickets for the same server.
This video shows how to set up a shell script to accept a positional parameter when called, pass that to a SQL script, accept the output from the statement back and then manipulate it in the Shell.
Via a live example, show how to extract information from SQL Server on Database, Connection and Server properties

697 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question