tokenizers v1: encode, decode and scaling, measured
Hugging Face's tokenizers v1 release candidate reworks the encode path with a hand-written SIMD splitter, word cache, allocation-free merge loop and native parallelism, delivering 3–30x faster encoding than v0.23 while producing identical token IDs.