Identify
Somebody sends you JBSWY3DPEBLW64TMMQ======. What is it?
import { identify } from "@agntn/encodings";
identify("JBSWY3DPEBLW64TMMQ======");
// [
// { encoding: "base32", confidence: 0.593, text: "Hello World",
// reasons: ["padding fits the block length", "decodes to readable text"], … },
// { encoding: "base32-crockford", confidence: 0.043, … },
// …
// ]
identify decodes the text in every registered encoding and keeps the ones that work. Then it ranks them by evidence, and the evidence is written down in reasons:
- A matching checksum. Base58Check and bech32 can't match by accident, so this outweighs everything else.
- Framing only one encoding writes. Padding that fits,
<~and~>, abegin 644line,=XXescapes, a0xprefix. - Decoded bytes that read as text. Valid UTF-8, mostly printable.
- A small alphabet. If the text fits in sixteen characters, hex is likelier than base64.
It leaves out anything that decodes to nothing, and anything that decodes to the input itself. Plain ASCII is technically valid Quoted-Printable. That's not an answer.
The number is a ranking
confidence runs from 0 to 1 and it sorts candidates. It isn't a probability. A base64 string of random bytes scores low, because there's no text and no checksum to vouch for it, and it's still the right answer. Read the reasons. Then decide.
Two encodings can write the same text. aGVsbG8gd29ybGQh is valid base64 and valid base64url, with the same score. The registry order breaks the tie, so base64 comes first. Ties are honest, not bugs.
Long text
Base58 decoding is quadratic in the length. Real base58 strings are short, addresses and keys and CIDs, so identify skips the base58 family above 1024 characters. The tools cap base58 at 10000 characters either way.
Only some encodings
identify(text, { encodings: ["base64", "base64url"], limit: 1 });
Your own registered encodings join the line-up without extra wiring.