Solved

Regular Expressions: Match the base letter of a unicode string

Posted on 2009-03-30
4
732 Views
Last Modified: 2013-12-17
Hello experts,

Is it possible to match the base letter of a unicode string?  If so, how do I do it?  So, for example, I have the word "hen" that I am looking for.  In my text file, I could have "hen" (which will match) and I could have "heñ" (which currently does not match).  I would like my regular expression or method thereof to be able to match both words.

So, Is there a regex tactic of which I am not aware that will match the base letter "n" when it comes across the unicode character ñ (and so on for every base letter)?

Thanks for shedding the light.
0
Comment
Question by:Gewgala
  • 2
  • 2
4 Comments
 
LVL 3

Expert Comment

by:gmrsecs
Comment Utility
basically, if you succeed to normalize your string in canonical mode, but I don't know how to do it in .net, you can use a simple reg exp like :
1)    he\u006E\p{M}*

where \u006E is the 'n' representation in unicode, and \p{M}* 0 or more diacritic signs. so this reg exp will match 'hen', but also heX(where X is a composition between \006E and    a diacritic(eg. \u0301))


anyway, the problem remains the canonical decomposition.
0
 
LVL 3

Accepted Solution

by:
gmrsecs earned 500 total points
Comment Utility
I've made some research and I saw that .net string object has Normalize method and you can transform your string before applying reg exp like:

s.Normalize(NormalizationForm.FormD)

It should work(but it is not tested).
0
 
LVL 7

Author Comment

by:Gewgala
Comment Utility
Thank you gmrsecs, that's exactly what I needed.  I applyed the NormalizationForm.FormD to my string, but I then ran a regex after that on the same string that stripped out all diacritic symbols.  So, for example, I ran this:

string s = <contents of file>;
string decoded = s.Normalize(Normalization.FormD);

Regex r = new Regex("\p{M}+", RegexOptions.Compiled);
decoded = r.Replace(decoded, "");

the string variable "decoded" would now contain the exact same content of the string variable "s" except all diacritic symbols would be stripped out, such as all ñ characters become simply n and so on, which I am them able to perform my matches on the decoded string and grab everything that I need.

Thanks!
0
 
LVL 7

Author Closing Comment

by:Gewgala
Comment Utility
Thanks again!
0

Featured Post

Threat Intelligence Starter Resources

Integrating threat intelligence can be challenging, and not all companies are ready. These resources can help you build awareness and prepare for defense.

Join & Write a Comment

Suggested Solutions

Title # Comments Views Activity
C# Export DataGridView 4 37
Code enhancement 5 12
dynamic menu in asp.net c# 11 25
Make a border less form movable 2 11
This article describes relatively difficult and non-obvious issues that are likely to arise when creating COM class in Visual Studio and deploying it by professional MSI-authoring tools. It is assumed that the reader is already familiar with the cla…
A long time ago (May 2011), I have written an article showing you how to create a DLL using Visual Studio 2005 to be hosted in SQL Server 2005. That was valid at that time and it is still valid if you are still using these versions. You can still re…
Learn how to match and substitute tagged data using PHP regular expressions. Demonstrated on Windows 7, but also applies to other operating systems. Demonstrated technique applies to PHP (all versions) and Firefox, but very similar techniques will w…
Explain concepts important to validation of email addresses with regular expressions. Applies to most languages/tools that uses regular expressions. Consider email address RFCs: Look at HTML5 form input element (with type=email) regex pattern: T…

772 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

10 Experts available now in Live!

Get 1:1 Help Now