Solved

Getting most common words from array

Posted on 2006-07-01
8
290 Views
Last Modified: 2012-06-21
I need a php script to find repeat words and output it.

Example text submitted from a textarea box:

drink energy
juice drink
hot cakes
fruit baskets
peanut butter and jelly
energy juice
fine wine
berry jelly

should return:

drink energy
juice drink
peanut butter and jelly
energy juice
berry jelly

because energy, juice, jelly and drink were most common, it should return those values from an array and also count it.


0
Comment
Question by:ray-solomon
  • 4
  • 4
8 Comments
 
LVL 29

Expert Comment

by:TeRReF
ID: 17025935
Something like this should work:
<?php

        $words = array('drink energy', 'juice drink', 'hot cakes', 'fruit baskets', 'peanut butter and jelly', 'energy juice', 'fine wine', 'berry jelly');
        $s = implode(' ', $words);
        $singlewords = array_unique(explode(' ', $s));
        //print_r($singlewords);
        //print($s);
        foreach($singlewords as $word) {
                preg_match_all('/'.$word.'/i', $s, $matches);
                $wordcount[$word] = count($matches[0]);
        }
        arsort($wordcount);
        $final_array = array();
        foreach($wordcount as $word=>$count) {
                foreach($words as $match) {
                        if (stripos($match, $word) !== false && !in_array($match, $final_array))
                                $final_array[] = $match;
                }
        }
        print_r($final_array);


?>
0
 
LVL 10

Author Comment

by:ray-solomon
ID: 17029212
Thanks TeRRef, but I get this error message:
Fatal error: Call to undefined function: stripos() in /home/...
0
 
LVL 29

Expert Comment

by:TeRReF
ID: 17029227
CHange this line:
                      if (stripos($match, $word) !== false && !in_array($match, $final_array))
to
                      if (strpos(strtolower($match), strtolower($word)) !== false && !in_array($match, $final_array))
0
 
LVL 10

Author Comment

by:ray-solomon
ID: 17031869
Here is what the array contains:

Array ( [0] => drink energy [1] => juice drink [2] => peanut butter and jelly [3] => berry jelly [4] => energy juice [5] => fine wine [6] => fruit baskets [7] => hot cakes )


it should look like this:

Array ( [0] => drink energy [1] => juice drink [2] => peanut butter and jelly [3] => berry jelly [4] => energy juice )


because:
fine wine, fruit baskets and hot cakes do not contain any words that have been repeated two or more times in the array.

Hope that makes sense. BTW, thanks for helping me so far.
0
How to improve team productivity

Quip adds documents, spreadsheets, and tasklists to your Slack experience
- Elevate ideas to Quip docs
- Share Quip docs in Slack
- Get notified of changes to your docs
- Available on iOS/Android/Desktop/Web
- Online/Offline

 
LVL 10

Author Comment

by:ray-solomon
ID: 17061204
Is there a way to make it output the most common words like I showed in my original question?
0
 
LVL 29

Accepted Solution

by:
TeRReF earned 500 total points
ID: 17061358
Sure. Sorry, I overlooked your last comment.
Here you go:

<?php

        $words = array('drink energy', 'juice drink', 'hot cakes', 'fruit baskets', 'peanut butter and jelly', 'energy juice', 'fine wine', 'berry jelly');
        $s = implode(' ', $words);
        $singlewords = array_unique(explode(' ', $s));
        //print_r($singlewords);
        //print($s);
        foreach ($singlewords as $word) {
                preg_match_all('/'.$word.'/i', $s, $matches);
                if (count($matches[0]) > 1)
                        $wordcount[$word] = count($matches[0]);
        }
        arsort($wordcount);
        $final_array = array();
        foreach ($wordcount as $word=>$count) {
                foreach ($words as $match) {
                        if (strpos(strtolower($match), strtolower($word)) !== false && !in_array($match, $final_array))
                                $final_array[] = $match;
                }
        }
        print_r($final_array);


?>

0
 
LVL 10

Author Comment

by:ray-solomon
ID: 17062135
Thank you! Awsome.
0
 
LVL 29

Expert Comment

by:TeRReF
ID: 17062314
You're welcome.
0

Featured Post

How your wiki can always stay up-to-date

Quip doubles as a “living” wiki and a project management tool that evolves with your organization. As you finish projects in Quip, the work remains, easily accessible to all team members, new and old.
- Increase transparency
- Onboard new hires faster
- Access from mobile/offline

Join & Write a Comment

Deprecated and Headed for the Dustbin By now, you have probably heard that some PHP features, while convenient, can also cause PHP security problems.  This article discusses one of those, called register_globals.  It is a thing you do not want.  …
Author Note: Since this E-E article was originally written, years ago, formal testing has come into common use in the world of PHP.  PHPUnit (http://en.wikipedia.org/wiki/PHPUnit) and similar technologies have enjoyed wide adoption, making it possib…
The viewer will learn how to count occurrences of each item in an array.
The viewer will learn how to create and use a small PHP class to apply a watermark to an image. This video shows the viewer the setup for the PHP watermark as well as important coding language. Continue to Part 2 to learn the core code used in creat…

706 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

22 Experts available now in Live!

Get 1:1 Help Now