Solved

REGEX Help for Domain Name

Posted on 2014-03-06
3
390 Views
Last Modified: 2014-03-07
Hello,

I am sifting through a page of text and want to use preg_match on it to find any Twitter links and capture them.

What is the regex I would use to find "http://twitter.com/anyuser" with "anyuser" being any string? The "http" should also look for "https"



Thank you!
0
Comment
Question by:EffinGood
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
3 Comments
 
LVL 6

Assisted Solution

by:Tony O'Byrne
Tony O'Byrne earned 400 total points
ID: 39911469
For the most part, the regex should be relatively straight-forward...

I'll break it up into two parts - the protocol and domain (http://twitter.com/ or https://www.twitter.com/), and the "anyuser" part...

https?://(?:www\.)?twitter\.com/

That matches the following:
http://twitter.com/
https://twitter.com/
http://www.twitter.com/
https://www.twitter.com/

If you need to escape the forward slashes, just add the backslash in front of each:
https?:\/\/(?:www\.)?twitter\.com\/

A note on the (?:www\.) part -
(?:) is a "non-capturing group".  Because the parenthesis capture a match for a backreference, it's handy to use the (?:) if you don't intend to backreference it.  However, if you don't care about backreferences at all, then you can remove the "?:" part (but leave the parenthesis.

On the "anyuser" part...

I'm treating this separately because I'm not entirely sure what characters are allowed in a twitter username.

A good place to start is:
[a-zA-Z]+

This would match all alpha characters as long as there are one or more.  However, twitter probably allows numbers, too:
[a-zA-Z0-9]+

... and underscores?
[a-zA-Z0-9_]+

... and hyphens?
[a-zA-Z0-9_-]+

So the entire regex so far is:
https?://(?:www\.)?twitter\.com/[a-zA-Z0-9_-]+

This is a pretty good starting-point.  If there are more characters allowed in twitter usernames, just add them before the ']'.  Be careful, though...  Some characters such as '?' and '#' mean something in the URL query-string, so they are probably not valid in usernames (though I don't know that for a fact.)

Hope this helps!

All the best,
Tony.
0
 
LVL 110

Accepted Solution

by:
Ray Paseur earned 100 total points
ID: 39912285
Please see this article.  It shows the thought process used to develop the REGEX to find a domain name.
http://www.experts-exchange.com/Web_Development/Web_Languages-Standards/PHP/A_7830-A-Quick-Tour-of-Test-Driven-Development.html

A Google search for PHP regular expression will turn up a lot of good learning resources.  This one is particularly on point:
http://www.php.net/manual/en/reference.pcre.pattern.syntax.php

Also, plan on giving yourself plenty of time to learn and experiment (as shown in the Test-Driven-Development article).  You're wading into an amazingly complex backwater of computer science, where the entire language is written in punctuation!
https://xkcd.com/208/
http://xkcd.com/1171/

If you decide you want a simpler approach, please post an example of the source document and I'll be glad to show you how to parse it with simple PHP statements.
0
 

Author Closing Comment

by:EffinGood
ID: 39913295
Thanks guys!

Tony's worked, and Ray, your article is awesome.
0

Featured Post

Independent Software Vendors: We Want Your Opinion

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Generating table dynamically is the most common issue faced by php developers.... So it seems there is a need of an article that explains the basic concept of generating tables dynamically. It just requires a basic knowledge of html and little maths…
This article discusses four methods for overlaying images in a container on a web page
The viewer will learn how to create and use a small PHP class to apply a watermark to an image. This video shows the viewer the setup for the PHP watermark as well as important coding language. Continue to Part 2 to learn the core code used in creat…
The viewer will learn how to create a basic form using some HTML5 and PHP for later processing. Set up your basic HTML file. Open your form tag and set the method and action attributes.: (CODE) Set up your first few inputs one for the name and …

730 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question