?
Solved

REGEX Help for Domain Name

Posted on 2014-03-06
3
Medium Priority
?
392 Views
Last Modified: 2014-03-07
Hello,

I am sifting through a page of text and want to use preg_match on it to find any Twitter links and capture them.

What is the regex I would use to find "http://twitter.com/anyuser" with "anyuser" being any string? The "http" should also look for "https"



Thank you!
0
Comment
Question by:EffinGood
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
3 Comments
 
LVL 6

Assisted Solution

by:Tony O'Byrne
Tony O'Byrne earned 1600 total points
ID: 39911469
For the most part, the regex should be relatively straight-forward...

I'll break it up into two parts - the protocol and domain (http://twitter.com/ or https://www.twitter.com/), and the "anyuser" part...

https?://(?:www\.)?twitter\.com/

That matches the following:
http://twitter.com/
https://twitter.com/
http://www.twitter.com/
https://www.twitter.com/

If you need to escape the forward slashes, just add the backslash in front of each:
https?:\/\/(?:www\.)?twitter\.com\/

A note on the (?:www\.) part -
(?:) is a "non-capturing group".  Because the parenthesis capture a match for a backreference, it's handy to use the (?:) if you don't intend to backreference it.  However, if you don't care about backreferences at all, then you can remove the "?:" part (but leave the parenthesis.

On the "anyuser" part...

I'm treating this separately because I'm not entirely sure what characters are allowed in a twitter username.

A good place to start is:
[a-zA-Z]+

This would match all alpha characters as long as there are one or more.  However, twitter probably allows numbers, too:
[a-zA-Z0-9]+

... and underscores?
[a-zA-Z0-9_]+

... and hyphens?
[a-zA-Z0-9_-]+

So the entire regex so far is:
https?://(?:www\.)?twitter\.com/[a-zA-Z0-9_-]+

This is a pretty good starting-point.  If there are more characters allowed in twitter usernames, just add them before the ']'.  Be careful, though...  Some characters such as '?' and '#' mean something in the URL query-string, so they are probably not valid in usernames (though I don't know that for a fact.)

Hope this helps!

All the best,
Tony.
0
 
LVL 111

Accepted Solution

by:
Ray Paseur earned 400 total points
ID: 39912285
Please see this article.  It shows the thought process used to develop the REGEX to find a domain name.
http://www.experts-exchange.com/Web_Development/Web_Languages-Standards/PHP/A_7830-A-Quick-Tour-of-Test-Driven-Development.html

A Google search for PHP regular expression will turn up a lot of good learning resources.  This one is particularly on point:
http://www.php.net/manual/en/reference.pcre.pattern.syntax.php

Also, plan on giving yourself plenty of time to learn and experiment (as shown in the Test-Driven-Development article).  You're wading into an amazingly complex backwater of computer science, where the entire language is written in punctuation!
https://xkcd.com/208/
http://xkcd.com/1171/

If you decide you want a simpler approach, please post an example of the source document and I'll be glad to show you how to parse it with simple PHP statements.
0
 

Author Closing Comment

by:EffinGood
ID: 39913295
Thanks guys!

Tony's worked, and Ray, your article is awesome.
0

Featured Post

On Demand Webinar: Networking for the Cloud Era

Ready to improve network connectivity? Watch this webinar to learn how SD-WANs and a one-click instant connect tool can boost provisions, deployment, and management of your cloud connection.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Things That Drive Us Nuts Have you noticed the use of the reCaptcha feature at EE and other web sites?  It wants you to read and retype something that looks like this. Insanity!  It's not EE's fault - that's just the way reCaptcha works.  But it i…
This article discusses four methods for overlaying images in a container on a web page
The viewer will learn how to count occurrences of each item in an array.
The viewer will learn how to create a basic form using some HTML5 and PHP for later processing. Set up your basic HTML file. Open your form tag and set the method and action attributes.: (CODE) Set up your first few inputs one for the name and …
Suggested Courses

764 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question