Solved

Regular Expressions: Match the base letter of a unicode string

Posted on 2009-03-30
4
753 Views
Last Modified: 2013-12-17
Hello experts,

Is it possible to match the base letter of a unicode string?  If so, how do I do it?  So, for example, I have the word "hen" that I am looking for.  In my text file, I could have "hen" (which will match) and I could have "heñ" (which currently does not match).  I would like my regular expression or method thereof to be able to match both words.

So, Is there a regex tactic of which I am not aware that will match the base letter "n" when it comes across the unicode character ñ (and so on for every base letter)?

Thanks for shedding the light.
0
Comment
Question by:Gewgala
  • 2
  • 2
4 Comments
 
LVL 3

Expert Comment

by:gmrsecs
ID: 24027083
basically, if you succeed to normalize your string in canonical mode, but I don't know how to do it in .net, you can use a simple reg exp like :
1)    he\u006E\p{M}*

where \u006E is the 'n' representation in unicode, and \p{M}* 0 or more diacritic signs. so this reg exp will match 'hen', but also heX(where X is a composition between \006E and    a diacritic(eg. \u0301))


anyway, the problem remains the canonical decomposition.
0
 
LVL 3

Accepted Solution

by:
gmrsecs earned 500 total points
ID: 24027179
I've made some research and I saw that .net string object has Normalize method and you can transform your string before applying reg exp like:

s.Normalize(NormalizationForm.FormD)

It should work(but it is not tested).
0
 
LVL 7

Author Comment

by:Gewgala
ID: 24068655
Thank you gmrsecs, that's exactly what I needed.  I applyed the NormalizationForm.FormD to my string, but I then ran a regex after that on the same string that stripped out all diacritic symbols.  So, for example, I ran this:

string s = <contents of file>;
string decoded = s.Normalize(Normalization.FormD);

Regex r = new Regex("\p{M}+", RegexOptions.Compiled);
decoded = r.Replace(decoded, "");

the string variable "decoded" would now contain the exact same content of the string variable "s" except all diacritic symbols would be stripped out, such as all ñ characters become simply n and so on, which I am them able to perform my matches on the decoded string and grab everything that I need.

Thanks!
0
 
LVL 7

Author Closing Comment

by:Gewgala
ID: 31564612
Thanks again!
0

Featured Post

Technology Partners: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Whatever be the reason, if you are working on web development side,  you will need day-today validation codes like email validation, date validation , IP address validation, phone validation on any of the edit page or say at the time of registration…
Introduction Hi all and welcome to my first article on Experts Exchange. A while ago, someone asked me if i could do some tutorials on object oriented programming. I decided to do them on C#. Now you may ask me, why's that? Well, one of the re…
Learn how to match and substitute tagged data using PHP regular expressions. Demonstrated on Windows 7, but also applies to other operating systems. Demonstrated technique applies to PHP (all versions) and Firefox, but very similar techniques will w…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…

685 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question