Base256
Base256
One emoji a byte, or any table of symbols a puzzle draws.
- Overhead
- +0% on random bytes
- Padding
- none
- Standard
- multiformats base256emoji
- Used by
- multiformats, multibase
Sample
- Out
- ๐ดโ๐๐๐๐ โ๐๐ฅบ๐๐ณ๐
- Back
- "hello world!"
Options
Access
- Import
import { base256 } from "@agntn/encodings/base256" - CLI
encodings encode base256 'hello world!' - Tryplayground with the sample above
One byte, one emoji. That's the whole trick. The emoji alphabet is the base256emoji table from multiformats, and it's the default. The roster puts it at +0%, since it counts characters. Don't pick it to save space though. In UTF-8 most emoji take four bytes, so the text runs about four times the data.
encode("base256", "gsmg", { multibase: true }); // "๐๐๐๐ท๐"
decode("base256", "๐๐๐๐ท๐", { multibase: true }).bytes; // "gsmg" as bytes
The rocket
Multibase puts ๐ in front of base256emoji text. Fair enough. But ๐ is also the table's first emoji, so it means byte 0 too. Read that string without multibase and you get a zero byte before gsmg. Nothing here strips it on a hunch, because a real zero byte looks exactly the same. identify tries both readings and puts the prefixed one first, since only that one reads as text.
Decoding skips ASCII whitespace and the emoji presentation selector. Copy โ๏ธ from a web page and it reads as the โ in the table.
Your own table
Puzzles keep inventing alphabets. Card suits, runes, a grid of emoji in a picture. symbols takes the table in order, digit 0 first, and spaces between symbols don't count.
decode("base256", "โฃ โ โฆ", { symbols: "โ โฅ โฆ โฃ" }).bytes; // [3, 0, 2]
encode("base256", new Uint8Array([3, 0]), { symbols: "โ โฅโฆโฃ" }); // "โฃโ "
Fewer than 256 symbols? Then you get digits, not bytes. Four suits read 0 to 3, and turning those digits into a message is the next layer of the puzzle. A byte past the end of the table throws on encode.
A symbol is what you see. A flag, a family emoji joined with ZWJ, a suit with or without its selector, each is one symbol. Zero-width characters are the exception, each one is a symbol of its own, so text that looks empty can still hold a table. Put ๐ต, ๐ฑ and ๐ต๐ฑ in one table and encode won't write ๐ต before ๐ฑ. That reads back as the flag. Better an error than other bytes, right?
Order from a sample
Sometimes the puzzle never lists the alphabet. It just shows it, and the order things first appear in is the order. sample reads the table that way, so repeats are fine there.
decode("base256", "แฆแ แข", { sample: "แ แขแ แฆแข" }).bytes; // [2, 0, 1]
Pass the ciphertext itself as the sample and every symbol gets numbered by its first sighting. Pass symbols or sample, not both.