Solved

RegEx issue

Posted on 2004-08-13
6
180 Views
Last Modified: 2010-04-15
ok, so I have a class which I've created that inherits from the class Regex [code shown below] and so far the expression has worked well for most cases except one.

Heres what its supposed to do:  Looks through any input tag (eg. <input type="text"...>) and find the attribute I specify followed by its value which I would like stored in the submatch.  In cases where I have a quoted value (tpe="text"), everything works fine... where it flunks is when the attribute has no quotes (sigle or double quotes, eg. type=TEXT).

Can anyone see what I'm doing wrong here?

protected class AttributeExpression : Regex
{
private static readonly RegexOptions PreDefOptions = RegexOptions.IgnoreCase | RegexOptions.Singleline | RegexOptions.Compiled;

public AttributeExpression(string attributeName) : base(" " + attributeName + "=\"([^\"]+)\"| " + attributeName + "='([^']+)'| " + attributeName + "=([^\\s]+)[\\s]", PreDefOptions){}
}
0
Comment
Question by:yleviel
  • 2
  • 2
  • 2
6 Comments
 
LVL 2

Expert Comment

by:davidastle
Comment Utility
Your problem is with the last section,
attributeName + "=([^\\s]+)[\\s]",

What your first group in this snippet, ([^\\s]+), does is seach for one or more non white space characters.  Therefore, it will only stop when you get to a white space.  After that, it tries to match [\\s], which is also looking for a non white space.  So you get to a white space, and try to match it with a non white space, and your match fails.
0
 
LVL 19

Accepted Solution

by:
drichards earned 120 total points
Comment Utility
No, the final [\s] looks FOR whitespace.

When I test your last expression it works - almost.  If the input looks like this:

    <input text=TEXT>

then the '>' is included in the capture.  I changed to this:

<att name>=([^>\s]+)[>\s]

Seems to work.  What was your test text?  Mine was dirt simple:

"<html><head><title>My Doc</title></head><body><form><input text=myText></input><input text='Some Text'></input></form></body></html>"

Your class with my small mod picked out both text= attributes and correctly captured "myText" and "Some Text".  I also found it a bit easier to name the groups:
----------------------------------------
protected class AttributeExpression : Regex
{
private static readonly RegexOptions PreDefOptions = RegexOptions.IgnoreCase | RegexOptions.Singleline | RegexOptions.Compiled;

public AttributeExpression(string attributeName) : base(" " + attributeName + "=\"(?<val>[^\"]+)\"| " + attributeName + "='(?<val>[^']+)'| " + attributeName + "=(?<val>[^>\\s]+)[>\\s]", PreDefOptions){}
}
0
 
LVL 2

Expert Comment

by:davidastle
Comment Utility
Oops, sorry
0
Better Security Awareness With Threat Intelligence

See how one of the leading financial services organizations uses Recorded Future as part of a holistic threat intelligence program to promote security awareness and proactively and efficiently identify threats.

 
LVL 2

Author Comment

by:yleviel
Comment Utility
drichards,

The regex you made works almost flawlessly... the only case where I've had problems is when the case of value="" shows up.  When this happens, the submatch returns "\"\"" (as in two quote symbols).  Anything you can think of to remedy this issue?

Thanks!
0
 
LVL 2

Author Comment

by:yleviel
Comment Utility
ok, I changed the expression to handle zero or more chars in the quotes. and this has worked for all my test cases.  If you see nothing wrong with the new expression I'll award you the points.

public AttributeExpression(string attributeName) : base(" " + attributeName + "=\"(?<val>[^\"]*)\"| " + attributeName + "='(?<val>[^']+)'| " + attributeName + "=(?<val>[^>\\s]+)[>\\s]", PreDefOptions){}
0
 
LVL 19

Expert Comment

by:drichards
Comment Utility
Just that you'll probably want the same change in the single quote expression (zero or more instead or 1 or more) and make sure there are no other terminal cases in the no-quote expression (anything other than whitespace and '>' that would end the match?).
0

Featured Post

Top 6 Sources for Identifying Threat Actor TTPs

Understanding your enemy is essential. These six sources will help you identify the most popular threat actor tactics, techniques, and procedures (TTPs).

Join & Write a Comment

Introduction This article series is supposed to shed some light on the use of IDisposable and objects that inherit from it. In essence, a more apt title for this article would be: using (IDisposable) {}. I’m just not sure how many people would ge…
Summary: Persistence is the capability of an application to store the state of objects and recover it when necessary. This article compares the two common types of serialization in aspects of data access, readability, and runtime cost. A ready-to…
In this seventh video of the Xpdf series, we discuss and demonstrate the PDFfonts utility, which lists all the fonts used in a PDF file. It does this via a command line interface, making it suitable for use in programs, scripts, batch files — any pl…
Illustrator's Shape Builder tool will let you combine shapes visually and interactively. This video shows the Mac version, but the tool works the same way in Windows. To follow along with this video, you can draw your own shapes or download the file…

762 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

8 Experts available now in Live!

Get 1:1 Help Now