Solved

How to find the character encoding type of a file in Linux?

Posted on 2011-03-21
7
909 Views
Last Modified: 2012-05-11
I have a .txt file and I need to determine what character encoding it is using so I can then convert other files to match it.

If I run "file myfile.txt", I get this info:
       "Non-ISO extended-ASCII text, with very long lines"

I know the file is ANSI but I need to determine exactly what type of ANSI file so I can convert other files to match it.

When I check the filetypes available in "iconv", I find these possibilities. How do I determine which one is the exact match?

ANSI_X3.4-1968
ANSI_X3.4-1986
ANSI_X3.4
ANSI_X3.110-1983
ANSI_X3.110
ASCII
MS-ANSI
WINDOWS-31J
WINDOWS-874
WINDOWS-936
WINDOWS-1250
WINDOWS-1251
WINDOWS-1252
WINDOWS-1253
WINDOWS-1254
WINDOWS-1255
WINDOWS-1256
WINDOWS-1257
WINDOWS-1258


0
Comment
Question by:bearclaws75
[X]
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 5
  • 2
7 Comments
 
LVL 31

Expert Comment

by:farzanj
ID: 35185608
Just use the command
unix2dos filename


And it should convert it to the DOS format.
or sometimes called
ux2dos
0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185616
If you want to go the other way,

issue this command

dos2unix filename
0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185634
I think the character encoding is UTF-8
0
Technology Partners: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

 

Author Comment

by:bearclaws75
ID: 35185695
farzani - I know the file is not UTF-8 because if I run "file otherfile.txt" on a different file, the output is:
     "UTF-8 Unicode text, with very long lines, with CRLF line terminators"

Howver, I ran "unix2dos myfile.txt" and it converted the file:
     "unix2dos: converting file myfile.txt to DOS format ..."

...but if I run "file myfile.txt", I get the same info:
       "Non-ISO extended-ASCII text, with very long lines"

"unix2dos" is a good command-line utility but, ultimately, i need to determine the exact character encoding so I can update my php scripts to generate the proper file type.


0
 
LVL 31

Expert Comment

by:farzanj
ID: 35185797
Well, I see your point but you can still call this utility from within PHP.  In any case let me look into it
0
 
LVL 31

Accepted Solution

by:
farzanj earned 500 total points
ID: 35185912
Well, I think it is very simple.  Basically you are converting the new line characters, that is about all.  Rest the remaining are the ASCII codes for characters which are the same.

So you need to convert line feed (\n) to carriage return (\r) and line feed.  Use a simple regular expression to do that.

So you are changing \n  to \r\n
0
 

Author Closing Comment

by:bearclaws75
ID: 35217944
I found this command which did the trick:

sed 's/\r$//' winfile.txt > unixfile.txt

I still wasn't able to determine the *exact* file encoding but this produced the desired results.

Thanks for the help!
0

Featured Post

Technology Partners: We Want Your Opinion!

We value your feedback.

Take our survey and automatically be enter to win anyone of the following:
Yeti Cooler, Amazon eGift Card, and Movie eGift Card!

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

Using dates in 'DOS' batch files has always been tricky as it has no built in ways of extracting date information.  There are many tricks using string manipulation to pull out parts of the %date% variable or output of the date /t command but these r…
The purpose of this article is to demonstrate how we can upgrade Python from version 2.7.6 to Python 2.7.10 on the Linux Mint operating system. I am using an Oracle Virtual Box where I have installed Linux Mint operating system version 17.2. Once yo…
This is a high-level webinar that covers the history of enterprise open source database use. It addresses both the advantages companies see in using open source database technologies, as well as the fears and reservations they might have. In this…
NetCrunch network monitor is a highly extensive platform for network monitoring and alert generation. In this video you'll see a live demo of NetCrunch with most notable features explained in a walk-through manner. You'll also get to know the philos…

688 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question