Guide

Code pages

EBCDIC and Latin-1 and Windows-1252. Read bytes as text in another character set or turn mojibake back into what it was

Bytes in a different character set

Every encoding here writes bytes as text. A code page doesn't. It says which character each byte already is. Byte 0xC8 is È in Latin-1 and H on an IBM mainframe. Same byte, different letter.

ts
import { charsets, hex } from "@agntn/encodings";

charsets.toText(hex.decode("c8c5d3d3d640e6d6d9d3c4"), { codepage: "ibm037" });
// "HELLO WORLD"
charsets.fromText("Grüße", { codepage: "ibm273" });
// Uint8Array [0xc7, 0x99, 0xd0, 0xa1, 0x85]

toText reads bytes, one character each. fromText goes back. A character the page doesn't have throws a DecodeError with its index. No ?, no guess.

Which pages?

codepageWhat it is
ibm037EBCDIC for the US and Canada
ibm273EBCDIC for Germany and Austria, with ä, ö, ü and ß
ibm500EBCDIC, the international one
ibm1140, ibm1141037 and 273 with € on byte 0x9F, where ¤ used to be
latin1ISO 8859-1. Byte n is code point n
windows1252What a browser calls Latin-1. Smart quotes and € on 0x80 to 0x9F

The EBCDIC tables come from IBM's own mappings, checked against glibc iconv. Windows-1252 reads the five bytes Microsoft never assigned as the C1 controls, the way every browser does.

Mojibake

Mojibake is text shown in the wrong code page. So you undo it in two steps. Read the text in the page it was shown in, then read those bytes in the page they were written in.

ts
const bytes = charsets.fromText("ÎÈ,Îø%_", { codepage: "ibm1141" });
new TextDecoder().decode(bytes);
// "vtkvplm"

Plain ASCII letters, printed as EBCDIC 1141. That's how a puzzle hides a Beaufort ciphertext in plain sight. Looks like line noise, turns out to be lowercase letters. Which page was it? Nothing in the text tells you, so identify doesn't guess. Try the likely ones.

CLI and agents

shell
encodings convert c8c5d3d3d640e6d6d9d3c4 --from hex --to ibm037
# HELLO WORLD

encodings convert 'ÎÈ,Îø%_' --from ibm1141
# vtkvplm

--from turns the input into bytes, --to writes them back, utf8 by default. Both take a code page, utf8, hex, base64 or raw. Agents get the same thing as encodings_charset_convert, with text, from and to.

text
hex → 11 bytes → ibm037:
"HELLO WORLD"

The subpath is @agntn/encodings/charsets, if the tables are all you need.