• Status: Solved
  • Priority: Medium
  • Security: Public
  • Views: 269
  • Last Modified:

PHP html-File search

Hello,

I need an PHP Script which open all HTML files on the public_html directory ( including subdirectorys ) and Search for an "string" ( non case sensitiv ) and the "title" of the page ( every html page have an title like: <title>Home</title>. It schould count how often this "string" appears on this site and print the result as an Link sorted by the number of hits.

For example my webspace contains 3 HTML pages:
- index.htm
- misc.htm
- kontakt.htm
I search for the word "Images", it exist in the Page "index.htm" once, in the misc.htm 8 of times. The output of the script should be:
Miscellaneous (8)
-> <a href="misc.htm">View Site</a>
Home (1)
-> <a href="misc.htm">View Site</a>

Some Information:
- Webserver Apache 1.3.31
- PHP-version: 4.x
- complete directory for local search: /home5/f6721805/public_html/ ( only files of the "public_html" will be accessable on the internet )

I have never written an php programm ( only other languages like Simatic S7 for Siemens SPS, HTML/CSS, VBS) so i think you peoples are faster in writing it than me. I think I'm able to integrate the script into my HTML Files.

----------------------------------------------------------
Sorry for my bad english, if i have something written incomprehensible please tell me. I will try to explain it again/better.
0
Kakashi
Asked:
Kakashi
  • 3
  • 2
1 Solution
 
thecode101Commented:
Try this out, if nothing else it is a good start:

<?php
$dir = "";
$search = "";

 if ($handle = opendir($dir)) {
      while (false !== ($file = readdir($handle))) {
      $contents = "";
              if ($file != "." && $file != "..") {
                         $fp = fopen($dir."/".$file, "r");
                         $contents = fread($fp, filesize($dir."/".$file));
                         fclose($fp);
                         $split = explode ("<title>",$contents);
                         $split = explode ("</title>",$split[1]);
                         $title = $split[0];
                         $numOfOccurences = substr_count ($contents,$search);
                         echo $title."(".$numOfOccurences.")"."<a href='".$dir."/".$file."'>View Site</a><br>";
              }
      }
}
closedir($handle);
?>
0
 
SashoCommented:
Here is more code to look at for ideas. But please remember that thecode101 answered first and his code should work as well.
<?PHP

function recursive_listdir($base) {
   static $filelist = array();
   static $dirlist = array();

   if(is_dir($base)) {
       $dh = opendir($base);
       while (false !== ($dir = readdir($dh))) {
           if (is_dir($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $subbase = $base ."/". $dir;
               $dirlist[] = $subbase;
               $subdirlist = recursive_listdir($subbase);
           } elseif(is_file($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $filelist[] = $base ."/". $dir;
           }
       }
       closedir($dh);
   }
   $array['dirs'] = $dirlist;
   $array['files'] = $filelist;
   return $array;
 }


$directory_structure=recursive_listdir(".");
foreach ($directory_structure['files'] as $file){
      $handle = fopen($file, "r");
      $lines = fread($handle, filesize($file));
      fclose($handle);

      $count=preg_match_all("/hello/i",$lines,$matches);
      preg_match("/<title>(.*)<\/title>/i",$lines, $matches);
      $title = $matches[1];

      print("$title($count) <a href=\"$file\">View Site</a><br>");
}

?>
0
 
KakashiAuthor Commented:
@ Sahso

wow your code is good. but i need your help.

- first i get some warnings look at http://www.synapstix.de/suche.php.
- Next thing this script search every file how can i confine the search only to *.htm|*.html files ?

- First thing i have done is that i have change your result output:
  if($count != 0){
       print("$title($count) <a href=\"$file\">View Site</a><br>");
   }
0
What Kind of Coding Program is Right for You?

There are many ways to learn to code these days. From coding bootcamps like Flatiron School to online courses to totally free beginner resources. The best way to learn to code depends on many factors, but the most important one is you. See what course is best for you.

 
SashoCommented:
Here is the fix to only do htm and html files:
<?PHP

function recursive_listdir($base) {
   static $filelist = array();
   static $dirlist = array();

   if(is_dir($base)) {
       $dh = opendir($base);
       while (false !== ($dir = readdir($dh))) {
           if (is_dir($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $subbase = $base ."/". $dir;
               $dirlist[] = $subbase;
               $subdirlist = recursive_listdir($subbase);
           } elseif(is_file($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $filelist[] = $base ."/". $dir;
           }
       }
       closedir($dh);
   }
   $array['dirs'] = $dirlist;
   $array['files'] = $filelist;
   return $array;
 }


$directory_structure=recursive_listdir(".");
//print_r($directory_structure['files']);

foreach ($directory_structure['files'] as $file){


      if (preg_match("/(.*)\.htm[l]*$/",$file) != 0 ){
            $handle = fopen($file, "r");
            $lines = fread($handle, filesize($file));
            fclose($handle);

            $count=preg_match_all("/hello/i",$lines,$matches);
            preg_match("/<title>(.*)<\/title>/i",$lines, $matches);
            $title = $matches[1];

            print("$title($count) <a href=\"$file\">View Site</a><br>");
      }
}

?>
0
 
SashoCommented:
Try this version to see if your Warnings go away:
<?PHP

function recursive_listdir($base) {
   static $filelist = array();
   static $dirlist = array();

   if(is_dir($base)) {
       $dh = opendir($base);
       while (false !== ($dir = readdir($dh))) {
           if (is_dir($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $subbase = $base ."/". $dir;
               $dirlist[] = $subbase;
               $subdirlist = recursive_listdir($subbase);
           } elseif(is_file($base ."/". $dir) && $dir !== '.' && $dir !== '..') {
               $filelist[] = $base ."/". $dir;
           }
       }
       closedir($dh);
   }
   $array['dirs'] = $dirlist;
   $array['files'] = $filelist;
   return $array;
 }


$directory_structure=recursive_listdir(".");
//print_r($directory_structure['files']);

foreach ($directory_structure['files'] as $file){


      if (preg_match("/(.*)\.htm[l]*$/",$file) != 0 ){
            $count = 0;
            $lines='';
            $handle = fopen($file, "r");
            if (filesize($file)!=0){
                  $lines = fread($handle, filesize($file));
            }
            fclose($handle);

            $count=preg_match_all("/hello/i",$lines,$matches);
            preg_match("/<title>(.*)<\/title>/i",$lines, $matches);
            $title = $matches[1];

            if($count != 0){
                   print("$title($count) <a href=\"$file\">View Site</a><br>");
               }
      }
}

?>
0
 
KakashiAuthor Commented:
Sasho Thanks a lot, you get the points ^^
0
Question has a verified solution.

Are you are experiencing a similar issue? Get a personalized answer when you ask a related question.

Have a better answer? Share it in a comment.

Join & Write a Comment

Featured Post

Free Tool: Path Explorer

An intuitive utility to help find the CSS path to UI elements on a webpage. These paths are used frequently in a variety of front-end development and QA automation tasks.

One of a set of tools we're offering as a way of saying thank you for being a part of the community.

  • 3
  • 2
Tackle projects and never again get stuck behind a technical roadblock.
Join Now