Encodings

Base256

One symbol per byte. The multiformats base256emoji table by default or any table of 2 to 256 symbols you give it
IDbase25611 / 14base256 ยท 256 characters
Encoding / Base256

Base256

One emoji a byte, or any table of symbols a puzzle draws.

Overhead
+0% on random bytes
Padding
none
Standard
multiformats base256emoji
Used by
multiformats, multibase
family 1 / 14

Sample

Out
๐Ÿ˜ดโœ‹๐Ÿ€๐Ÿ€๐Ÿ˜“๐Ÿ˜…โœ”๐Ÿ˜“๐Ÿฅบ๐Ÿ€๐Ÿ˜ณ๐Ÿ‘
Back
"hello world!"

Options

alphabetstring
default emoji
emoji, the base256emoji table of multiformats
symbolsstring
optional
Symbols for 0 upward in place of the alphabet, 2 to 256 of them; fewer read as digits
samplestring
optional
Text whose symbols, in order of first appearance, make the table
multibaseboolean
default false
The ๐Ÿš€ multibase prefix of base256emoji

Access

Importimport { base256 } from "@agntn/encodings/base256"
CLIencodings encode base256 'hello world!'
Tryplayground with the sample above

One byte, one emoji. That's the whole trick. The emoji alphabet is the base256emoji table from multiformats, and it's the default. The roster puts it at +0%, since it counts characters. Don't pick it to save space though. In UTF-8 most emoji take four bytes, so the text runs about four times the data.

ts
encode("base256", "gsmg", { multibase: true });  // "๐Ÿš€๐Ÿ˜๐ŸŒˆ๐ŸŒท๐Ÿ˜"
decode("base256", "๐Ÿš€๐Ÿ˜๐ŸŒˆ๐ŸŒท๐Ÿ˜", { multibase: true }).bytes;  // "gsmg" as bytes

The rocket

Multibase puts ๐Ÿš€ in front of base256emoji text. Fair enough. But ๐Ÿš€ is also the table's first emoji, so it means byte 0 too. Read that string without multibase and you get a zero byte before gsmg. Nothing here strips it on a hunch, because a real zero byte looks exactly the same. identify tries both readings and puts the prefixed one first, since only that one reads as text.

Decoding skips ASCII whitespace and the emoji presentation selector. Copy โ˜„๏ธ from a web page and it reads as the โ˜„ in the table.

Your own table

Puzzles keep inventing alphabets. Card suits, runes, a grid of emoji in a picture. symbols takes the table in order, digit 0 first, and spaces between symbols don't count.

ts
decode("base256", "โ™ฃ โ™  โ™ฆ", { symbols: "โ™  โ™ฅ โ™ฆ โ™ฃ" }).bytes;  // [3, 0, 2]
encode("base256", new Uint8Array([3, 0]), { symbols: "โ™ โ™ฅโ™ฆโ™ฃ" });  // "โ™ฃโ™ "

Fewer than 256 symbols? Then you get digits, not bytes. Four suits read 0 to 3, and turning those digits into a message is the next layer of the puzzle. A byte past the end of the table throws on encode.

A symbol is what you see. A flag, a family emoji joined with ZWJ, a suit with or without its selector, each is one symbol. Zero-width characters are the exception, each one is a symbol of its own, so text that looks empty can still hold a table. Put ๐Ÿ‡ต, ๐Ÿ‡ฑ and ๐Ÿ‡ต๐Ÿ‡ฑ in one table and encode won't write ๐Ÿ‡ต before ๐Ÿ‡ฑ. That reads back as the flag. Better an error than other bytes, right?

Order from a sample

Sometimes the puzzle never lists the alphabet. It just shows it, and the order things first appear in is the order. sample reads the table that way, so repeats are fine there.

ts
decode("base256", "แšฆแš แšข", { sample: "แš แšขแš  แšฆแšข" }).bytes;  // [2, 0, 1]

Pass the ciphertext itself as the sample and every symbol gets numbered by its first sighting. Pass symbols or sample, not both.