Solved

Regular Expressions: Match the base letter of a unicode string

Posted on 2009-03-30
4
747 Views
Last Modified: 2013-12-17
Hello experts,

Is it possible to match the base letter of a unicode string?  If so, how do I do it?  So, for example, I have the word "hen" that I am looking for.  In my text file, I could have "hen" (which will match) and I could have "heñ" (which currently does not match).  I would like my regular expression or method thereof to be able to match both words.

So, Is there a regex tactic of which I am not aware that will match the base letter "n" when it comes across the unicode character ñ (and so on for every base letter)?

Thanks for shedding the light.
0
Comment
Question by:Gewgala
  • 2
  • 2
4 Comments
 
LVL 3

Expert Comment

by:gmrsecs
ID: 24027083
basically, if you succeed to normalize your string in canonical mode, but I don't know how to do it in .net, you can use a simple reg exp like :
1)    he\u006E\p{M}*

where \u006E is the 'n' representation in unicode, and \p{M}* 0 or more diacritic signs. so this reg exp will match 'hen', but also heX(where X is a composition between \006E and    a diacritic(eg. \u0301))


anyway, the problem remains the canonical decomposition.
0
 
LVL 3

Accepted Solution

by:
gmrsecs earned 500 total points
ID: 24027179
I've made some research and I saw that .net string object has Normalize method and you can transform your string before applying reg exp like:

s.Normalize(NormalizationForm.FormD)

It should work(but it is not tested).
0
 
LVL 7

Author Comment

by:Gewgala
ID: 24068655
Thank you gmrsecs, that's exactly what I needed.  I applyed the NormalizationForm.FormD to my string, but I then ran a regex after that on the same string that stripped out all diacritic symbols.  So, for example, I ran this:

string s = <contents of file>;
string decoded = s.Normalize(Normalization.FormD);

Regex r = new Regex("\p{M}+", RegexOptions.Compiled);
decoded = r.Replace(decoded, "");

the string variable "decoded" would now contain the exact same content of the string variable "s" except all diacritic symbols would be stripped out, such as all ñ characters become simply n and so on, which I am them able to perform my matches on the decoded string and grab everything that I need.

Thanks!
0
 
LVL 7

Author Closing Comment

by:Gewgala
ID: 31564612
Thanks again!
0

Featured Post

Master Your Team's Linux and Cloud Stack!

The average business loses $13.5M per year to ineffective training (per 1,000 employees). Keep ahead of the competition and combine in-person quality with online cost and flexibility by training with Linux Academy.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Many of us here at EE write code. Many of us write exceptional code; just as many of us write exception-prone code. As we all should know, exceptions are a mechanism for handling errors which are typically out of our control. From database errors, t…
A long time ago (May 2011), I have written an article showing you how to create a DLL using Visual Studio 2005 to be hosted in SQL Server 2005. That was valid at that time and it is still valid if you are still using these versions. You can still re…
Learn how to match and substitute tagged data using PHP regular expressions. Demonstrated on Windows 7, but also applies to other operating systems. Demonstrated technique applies to PHP (all versions) and Firefox, but very similar techniques will w…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…

809 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question