Solved

development site is indexed by google even though behind htpasswd

Posted on 2016-11-11
7
26 Views
Last Modified: 2016-11-15
As stated in the title. My development site is protected with an htpasswd, yet somehow google is indexing it. how is that possible?
0
Comment
Question by:jblayney
  • 3
  • 2
  • 2
7 Comments
 
LVL 12

Expert Comment

by:Phil Phillips
ID: 41884313
htpasswd should prevent Google from accessing the site, though you might want to double check that you have the config set up to lock down all URLs. In the web logs, do you still see Google accessing it with successful response codes?

Also, it could have just indexed a previous version. If that's the case, it'll take some time for it to age out.  To speed up the process, you can login to Google webmaster tools and remove the site from the index.
0
 
LVL 23

Expert Comment

by:Dr. Klahn
ID: 41884320
Update robots.txt to exclude all robots on that site:

User-agent: *
Disallow: /

Open in new window


I've noticed that Googlebot occasionally visits my sites with faked browser credentials to avoid complying with robots.txt, often enough that I wrote a mod_rewrite rule to block the behavior. So I'd also throw in an exclusion rule for the Googlebot IP ranges:

# Googlebot's various /24 blocks
66.249.64.0/22
66.249.69.0/24
66.249.73.0/24
66.249.79.0/24

Open in new window

0
 
LVL 1

Author Comment

by:jblayney
ID: 41886558
thanks for responding, you mean this?

Order Allow,Deny
Deny from 66.249.64.0/22
Deny from 66.249.69.0/24
Deny from 66.249.73.0/24
Deny from 66.249.79.0/24
Allow from all

Open in new window

0
Highfive Gives IT Their Time Back

Highfive is so simple that setting up every meeting room takes just minutes and every employee will be able to start or join a call from any room with ease. Never be called into a meeting just to get it started again. This is how video conferencing should work!

 
LVL 1

Author Comment

by:jblayney
ID: 41886559
Phil,

where do I do this?
In the web logs, do you still see Google accessing it with successful response codes?
0
 
LVL 12

Assisted Solution

by:Phil Phillips
Phil Phillips earned 250 total points
ID: 41886654
It depends how you have Apache configured to store your logs. If you're on Linux, a common default place is: /var/log/httpd or /var/log/apache2
0
 
LVL 23

Accepted Solution

by:
Dr. Klahn earned 250 total points
ID: 41887051
thanks for responding, you mean this?

(... list of IP blocks)

That's the one.  It can be done in the main config file or in an htaccess file.  If you don't want googlebot in the site at all, it is more efficient to do it in iptables which blocks the request before it gets to Apache.
0
 
LVL 1

Author Closing Comment

by:jblayney
ID: 41888001
thank you
0

Featured Post

Do You Know the 4 Main Threat Actor Types?

Do you know the main threat actor types? Most attackers fall into one of four categories, each with their own favored tactics, techniques, and procedures.

Join & Write a Comment

Introduction As you’re probably aware the HTTP protocol offers basic / weak authentication, which in combination with the relevant configuration on your web server, provides the ability to password protect all or part of your host.  If you were not…
Hi, in this article I'm going to teach you how to run your own site, and how to let people in (without IP). I'll talk about and explain each step... :) By the way, everything in this Tutorial is completely free and legal. This article is for …
It is a freely distributed piece of software for such tasks as photo retouching, image composition and image authoring. It works on many operating systems, in many languages.
Polish reports in Access so they look terrific. Take yourself to another level. Equations, Back Color, Alternate Back Color. Write easy VBA Code. Tighten space to use less pages. Launch report from a menu, considering criteria only when it is filled…

743 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

14 Experts available now in Live!

Get 1:1 Help Now