Solved

MS SQL for count of words in a series of sentences.

Posted on 2013-11-12
3
246 Views
Last Modified: 2013-11-13
Need a query that counts number of occurances of the word similar to count(field).  I have a db where each record has a sentence.  Count(sententFiled) returns the number of sentences with a word.  This is lower than the actual number of words because some words occur more than once in a sentence.

Thanks!
0
Comment
Question by:HyperBPP
  • 2
3 Comments
 
LVL 9

Expert Comment

by:QuinnDex
ID: 39642870
try this function

create function [dbo].[fnParseWords](@str varchar(max), @delimiter varchar(30)='%[^a-zA-Z0-9\_]%')
returns @result table(word varchar(max))
begin
    if left(@delimiter,1)<>'%' set @delimiter='%'+@delimiter;
    if right(@delimiter,1)<>'%' set @delimiter+='%';
    set @str=rtrim(@str);
    declare @pi int=PATINDEX(@delimiter,@str);

    while @pi>0 begin
        insert into @result select LEFT(@str,@pi-1) where @pi>1;
        set @str=RIGHT(@str,len(@str)-@pi);
        set @pi=PATINDEX(@delimiter,@str);
    end

    insert into @result select @str where LEN(@str)>0;
    return;
end
go

select COUNT(*)
from webqueries q
cross apply dbo.fnParseWords(cast(q.qQuestion as varchar(max)),default) pw
where pw.word not in ('and','is','a','the'/* plus whatever else you need to exclude */)

Open in new window

0
 
LVL 48

Accepted Solution

by:
PortletPaul earned 500 total points
ID: 39643700
this will count the number of times a string is found in a larger string:

( len(sentence) - len(replace(sentence,'word','') ) / len('word')

{+edit} and just sum that if aggregating
sum( ( len(sentence) - len(replace(sentence,'word','') ) / len('word') )
0
 
LVL 48

Expert Comment

by:PortletPaul
ID: 39643706
example, sample data: 5 sentences containing 'amet', 10 occurences of 'amet':
    CREATE TABLE SentenceTable
    	([SentenceField] varchar(800)) 
    ;
    	
    INSERT INTO SentenceTable
    	([SentenceField])
    VALUES
    	('Lorem ipsum dolor sit amet, consectetur adipiscing elit amet.'),
    	('Fusce euismod justo id rhoncus lobortis.'),
    	('Nunc rhoncus amet risus vitae metus amet laoreet placerat amet.'),
    	('Nam vel nunc dapibus, suscipit eros ut, imperdiet erat.'),
    	('Proin ut enim fringilla, iaculis erat nec, mattis risus.'),
    	('Donec ac leo egestas, amet euismod velit id, blandit nunc.'),
    	('Vivamus vestibulum est non purus faucibus mattis.'),
    	('Vivamus in tortor ultrices, cursus massa eget, ornare justo.'),
    	('Etiam lobortis nunc nec commodo pretium.'),
    	('Nam in neque et mauris lobortis euismod.'),
    	('Suspendisse amet eget neque malesuada, amet cursus ligula et, sodales odio amet.'),
    	('Ut eget est facilisis, aliquet leo ac, posuere enim.'),
    	('Praesent consequat augue sed erat fermentum, sit amet fringilla turpis malesuada.'),
    	('Curabitur eget eros eget massa tempor interdum.')
    ;

**Query 1**:

    declare @word as varchar(100)
    set @word = 'amet'
    
    SELECT
      count(*)
    , sum( (len(SentenceField) - len(replace(SentenceField,@word,''))) / len(@word) )
    FROM SentenceTable
    WHERE SentenceField LIKE '%' + @word + '%'
    

**[Results][2]**:
    
    | COLUMN_0 | COLUMN_1 |
    |----------|----------|
    |        5 |       10 |


**Query 2**:

    declare @word as varchar(100)
    set @word = 'amet'
    
    SELECT
      SentenceField
    , (len(SentenceField) - len(replace(SentenceField,@word,''))) / len(@word)
    FROM SentenceTable
    WHERE SentenceField LIKE '%' + @word + '%'
    

**[Results][3]**:
    
    |                                                                     SENTENCEFIELD | COLUMN_1 |
    |-----------------------------------------------------------------------------------|----------|
    |                     Lorem ipsum dolor sit amet, consectetur adipiscing elit amet. |        2 |
    |                   Nunc rhoncus amet risus vitae metus amet laoreet placerat amet. |        3 |
    |                        Donec ac leo egestas, amet euismod velit id, blandit nunc. |        1 |
    |  Suspendisse amet eget neque malesuada, amet cursus ligula et, sodales odio amet. |        3 |
    | Praesent consequat augue sed erat fermentum, sit amet fringilla turpis malesuada. |        1 |



  [1]: http://sqlfiddle.com/#!3/8dc0e/1

Open in new window

0

Featured Post

PRTG Network Monitor: Intuitive Network Monitoring

Network Monitoring is essential to ensure that computer systems and network devices are running. Use PRTG to monitor LANs, servers, websites, applications and devices, bandwidth, virtual environments, remote systems, IoT, and many more. PRTG is easy to set up & use.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

If you have heard of RFC822 date formats, they can be quite a challenge in SQL Server. RFC822 is an Internet standard format for email message headers, including all dates within those headers. The RFC822 protocols are available in detail at:   ht…
Ever wondered why sometimes your SQL Server is slow or unresponsive with connections spiking up but by the time you go in, all is well? The following article will show you how to install and configure a SQL job that will send you email alerts includ…
Using examples as well as descriptions, and references to Books Online, show the documentation available for date manipulation functions and by using a select few of these functions, show how date based data can be manipulated with these functions.
This video shows how to set up a shell script to accept a positional parameter when called, pass that to a SQL script, accept the output from the statement back and then manipulate it in the Shell.

813 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

19 Experts available now in Live!

Get 1:1 Help Now