Code pages
Bytes in a different character set
Every encoding here writes bytes as text. A code page doesn't. It says which character each byte already is. Byte 0xC8 is È in Latin-1 and H on an IBM mainframe. Same byte, different letter.
import { charsets, hex } from "@agntn/encodings";
charsets.toText(hex.decode("c8c5d3d3d640e6d6d9d3c4"), { codepage: "ibm037" });
// "HELLO WORLD"
charsets.fromText("Grüße", { codepage: "ibm273" });
// Uint8Array [0xc7, 0x99, 0xd0, 0xa1, 0x85]
toText reads bytes, one character each. fromText goes back. A character the page doesn't have throws a DecodeError with its index. No ?, no guess.
Which pages?
codepage | What it is |
|---|---|
ibm037 | EBCDIC for the US and Canada |
ibm273 | EBCDIC for Germany and Austria, with ä, ö, ü and ß |
ibm500 | EBCDIC, the international one |
ibm1140, ibm1141 | 037 and 273 with € on byte 0x9F, where ¤ used to be |
latin1 | ISO 8859-1. Byte n is code point n |
windows1252 | What a browser calls Latin-1. Smart quotes and € on 0x80 to 0x9F |
The EBCDIC tables come from IBM's own mappings, checked against glibc iconv. Windows-1252 reads the five bytes Microsoft never assigned as the C1 controls, the way every browser does.
Mojibake
Mojibake is text shown in the wrong code page. So you undo it in two steps. Read the text in the page it was shown in, then read those bytes in the page they were written in.
const bytes = charsets.fromText("ÎÈ,Îø%_", { codepage: "ibm1141" });
new TextDecoder().decode(bytes);
// "vtkvplm"
Plain ASCII letters, printed as EBCDIC 1141. That's how a puzzle hides a Beaufort ciphertext in plain sight. Looks like line noise, turns out to be lowercase letters. Which page was it? Nothing in the text tells you, so identify doesn't guess. Try the likely ones.
CLI and agents
encodings convert c8c5d3d3d640e6d6d9d3c4 --from hex --to ibm037
# HELLO WORLD
encodings convert 'ÎÈ,Îø%_' --from ibm1141
# vtkvplm
--from turns the input into bytes, --to writes them back, utf8 by default. Both take a code page, utf8, hex, base64 or raw. Agents get the same thing as encodings_charset_convert, with text, from and to.
hex → 11 bytes → ibm037:
"HELLO WORLD"
The subpath is @agntn/encodings/charsets, if the tables are all you need.