Solved

How to find the character encoding type of a file in Linux?

Posted on 2011-03-21
7
900 Views
Last Modified: 2012-05-11
I have a .txt file and I need to determine what character encoding it is using so I can then convert other files to match it.

If I run "file myfile.txt", I get this info:
       "Non-ISO extended-ASCII text, with very long lines"

I know the file is ANSI but I need to determine exactly what type of ANSI file so I can convert other files to match it.

When I check the filetypes available in "iconv", I find these possibilities. How do I determine which one is the exact match?

ANSI_X3.4-1968
ANSI_X3.4-1986
ANSI_X3.4
ANSI_X3.110-1983
ANSI_X3.110
ASCII
MS-ANSI
WINDOWS-31J
WINDOWS-874
WINDOWS-936
WINDOWS-1250
WINDOWS-1251
WINDOWS-1252
WINDOWS-1253
WINDOWS-1254
WINDOWS-1255
WINDOWS-1256
WINDOWS-1257
WINDOWS-1258


0
Comment
Question by:bearclaws75
  • 5
  • 2
7 Comments
 
LVL 31

Expert Comment

by:farzanj
ID: 35185608
Just use the command
unix2dos filename


And it should convert it to the DOS format.
or sometimes called
ux2dos
0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185616
If you want to go the other way,

issue this command

dos2unix filename
0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185634
I think the character encoding is UTF-8
0
Master Your Team's Linux and Cloud Stack!

The average business loses $13.5M per year to ineffective training (per 1,000 employees). Keep ahead of the competition and combine in-person quality with online cost and flexibility by training with Linux Academy.

 

Author Comment

by:bearclaws75
ID: 35185695
farzani - I know the file is not UTF-8 because if I run "file otherfile.txt" on a different file, the output is:
     "UTF-8 Unicode text, with very long lines, with CRLF line terminators"

Howver, I ran "unix2dos myfile.txt" and it converted the file:
     "unix2dos: converting file myfile.txt to DOS format ..."

...but if I run "file myfile.txt", I get the same info:
       "Non-ISO extended-ASCII text, with very long lines"

"unix2dos" is a good command-line utility but, ultimately, i need to determine the exact character encoding so I can update my php scripts to generate the proper file type.


0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185797
Well, I see your point but you can still call this utility from within PHP.  In any case let me look into it
0
 
LVL 31

Accepted Solution

by:
farzanj earned 500 total points
ID: 35185912
Well, I think it is very simple.  Basically you are converting the new line characters, that is about all.  Rest the remaining are the ASCII codes for characters which are the same.

So you need to convert line feed (\n) to carriage return (\r) and line feed.  Use a simple regular expression to do that.

So you are changing \n  to \r\n
0
 

Author Closing Comment

by:bearclaws75
ID: 35217944
I found this command which did the trick:

sed 's/\r$//' winfile.txt > unixfile.txt

I still wasn't able to determine the *exact* file encoding but this produced the desired results.

Thanks for the help!
0

Featured Post

Master Your Team's Linux and Cloud Stack

Come see why top tech companies like Mailchimp and Media Temple use Linux Academy to build their employee training programs.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Using dates in 'DOS' batch files has always been tricky as it has no built in ways of extracting date information.  There are many tricks using string manipulation to pull out parts of the %date% variable or output of the date /t command but these r…
The purpose of this article is to fix the unknown display problem in Linux Mint operating system. After installing the OS if you see Display monitor is not recognized then we can install "MESA" utilities to fix this problem or we can install additio…
Two types of users will appreciate AOMEI Backupper Pro: 1 - Those with PCIe drives (and haven't found cloning software that works on them). 2 - Those who want a fast clone of their boot drive (no re-boots needed) and it can clone your drive wh…

832 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question