Solved

Convert NOT IN Sub-Select to LEFT JOIN

Posted on 2011-09-30
12
318 Views
Last Modified: 2012-05-12
When I run this query:

select distinct idQuestion from Question
where parentid is null AND idQuestion not in (select distinct parentid from Question where parentid is not null)
order by idQuestion

Open in new window


The query never finishes after waiting several minutes. I've researched the issue and the conclusion is to use a LEFT JOIN instead of a sub-select. However, I'm not sure how to properly construct the query with these limitations.
0
Comment
Question by:BCRobert
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 5
  • 4
  • 3
12 Comments
 
LVL 143

Expert Comment

by:Guy Hengel [angelIII / a3]
ID: 36891743
please try this:

select q.idQuestion 
from Question q
where q.parentid is null 
AND NOT EXISTS ( select NULL from Question o where o.parentid = q.idQuestion )
order by idQuestion 

Open in new window


and please ensure you have an index on parentid
0
 
LVL 5

Expert Comment

by:eridanix
ID: 36891750
Hi,

I mean you have little nonsense in your query, becouse you compare idQuestion with parentid

The correct query should only be:
select distinct idQuestion
from Question
where parentid is not null
order by idQuestion

Maybe, you need get something else. So try to explain what shloud be result of this query.
0
 

Author Comment

by:BCRobert
ID: 36891794
@angellll: parentid isn't indexed because it's not unique. This query is taking just as long.

@eridanix: That query won't work. Here's what I'm trying to find:

Find the question IDs that are not a parent ID of another question and the parentid is null.

idQuestion    |    data     |    parentid
123               |    abc      |       NULL
321               |    def       |        123
444               |    qqq       |    NULL

Open in new window


The query should return just '444' as its parentid is null AND is not a parentid of another question.
0
The Ultimate Checklist to Optimize Your Website

Websites are getting bigger and complicated by the day. Video, images, custom fonts are all great for showcasing your product/service. But the price to pay in terms of reduced page load times and ultimately, decreased sales, can lead to some difficult decisions about what to cut.

 
LVL 5

Expert Comment

by:eridanix
ID: 36891882
select distinct idQuestion
from Question
where parentid is null and idQuestion NOT IN (select idQuestion from Question where parentid is not null)
order by idQuestion
0
 
LVL 143

Accepted Solution

by:
Guy Hengel [angelIII / a3] earned 500 total points
ID: 36891919
>parentid isn't indexed because it's not unique. T

indexes can be on non-unique fields.
please create such an index.
0
 
LVL 143

Expert Comment

by:Guy Hengel [angelIII / a3]
ID: 36891930
edit: you might not know the a primary key is a UNIQUE index, under the hood, but you can created non-unique indexes.
actually, a index is non-unique by default, unless you specify it to add the restriction that each value can only be present once.
0
 

Author Comment

by:BCRobert
ID: 36891937
@eridanix: That will not return the results I need. I need to make sure the the idQuestion IS NOT the parentid of another question.
0
 

Author Comment

by:BCRobert
ID: 36891968
@angellll: The parentid is already an index, apparently (I didn't create the database/tables):

parentIdIdx      BTREE      No      No      parentId      41002      A      YES      
0
 
LVL 5

Expert Comment

by:eridanix
ID: 36892014
select idQuestion
from dbo.Questions
where idQuestion IN (select idQuestion from Question where parentid is null) AND idQuestion NOT IN  (select parentid from Question where parentid is not null)
0
 

Author Comment

by:BCRobert
ID: 36892103
@angellll: It seems that re-creating the index helped tremendously. Not sure why, but the query finished after about 45 seconds, which is still slow, but it's not a query we'll be running often.
0
 

Author Closing Comment

by:BCRobert
ID: 36892108
Creating/recreating the index helped the initial query finished. Converting it to LEFT JOIN/NOT IN isn't needed.
0
 
LVL 143

Expert Comment

by:Guy Hengel [angelIII / a3]
ID: 36892237
please try to remove the 2 DISTINCT in the query.

IN ( SELECT DISTINCT  ... ) will be the same, but slower normally, as
IN ( SELECT ... )

the NOT EXISTS () version I posted should be fastest ...
0

Featured Post

Efficient way to get backups off site to Azure

This user guide provides instructions on how to deploy and configure both a StoneFly Scale Out NAS Enterprise Cloud Drive virtual machine and Veeam Cloud Connect in the Microsoft Azure Cloud.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Occasionally there is a need to clean table columns, especially if you have inherited legacy data. There are obviously many ways to accomplish that, including elaborate UPDATE queries with anywhere from one to numerous REPLACE functions (even within…
This post looks at MongoDB and MySQL, and covers high-level MongoDB strengths, weaknesses, features, and uses from the perspective of an SQL user.
NetCrunch network monitor is a highly extensive platform for network monitoring and alert generation. In this video you'll see a live demo of NetCrunch with most notable features explained in a walk-through manner. You'll also get to know the philos…
Do you want to know how to make a graph with Microsoft Access? First, create a query with the data for the chart. Then make a blank form and add a chart control. This video also shows how to change what data is displayed on the graph as well as form…

690 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question