Tokenizer Comparison Cheat Sheet
This cheat sheet compares the capabilities, API designs, and performance features of three major tokenizer libraries: SentencePiece (new pybind11 API), Hugging Face Tokenizers, and OpenAI's tiktoken.
Note: Please feel free to send a PR if you find any problems or inaccuracies.
———
1. Capability Matrix
Feature / CapabilitySentencePieceHugging Face tokenizerstiktokenCompared Version>=v0.2.2v0.23.1v0.13.0Native BackendC++ / pybind11Rust / PyO3Rust / PyO3Supported AlgorithmsBPE, Unigram, Char, WordBPE, WordPiece, Unigram, WordLevelBPEOOV / Unknown Handling<unk> token, byte_fallback<unk> token, byte_fallback (BPE/Unigram), or Byte-level BPE (no <unk>)Byte-level BPE (no <unk> token)Training Support✅ Yes (via SentencePieceTrainer)✅ Yes (via Trainer classes)❌ NoGIL Release (Parallelism)✅ Yes (Releases GIL in almost all cases including single/batch encode & decode; supports Free-Threading)⚠️ Partial (Releases GIL during core Rust execution, but string copying is done with GIL held)⚠️ Partial (Releases GIL during core Rust execution, but python-level threading and conversions hold GIL)Encode Offset Mapping (Unicode)✅ Yes (via return_type='offset_mapping')✅ Yes (via Encoding.offsets)❌ NoDecode Offset Mapping (Unicode)✅ Yes (via return_type='offset_mapping', includes text key)❌ No⚠️ Partial (via Python decode_with_offsets; start offsets only)Raw Byte Offset Mapping✅ Yes (triggered by bytes input or return_bytes=True for both encode and decode)❌ No❌ NoDirect UTF-8 Bytes I/O✅ Yes (accepts bytes for encode, returns bytes for decode)❌ No (requires str)⚠️ Partial (requires str for encode, returns bytes via decode_bytes)Zero-Copy Input (str/bytes)✅ Yes (zero-copy for both str and bytes via Pybind11 memory view)❌ No (always copies input str to Rust owned String via PyO3)⚠️ Partial (zero-copy for str under some conditions, no bytes support)NumPy Array Output (Encode)✅ Yes (Direct to NumPy, zero-copy)❌ No (Transformers wrapper only)✅ Yes (via encode_to_numpy, zero-copy)NumPy Array Input (Decode)✅ Yes (Accepts 1D/2D; zero-copy on best-effort basis)✅ Yes (Accepts 1D/2D; copied)✅ Yes (Accepts 1D/2D; copied)Batch Encoding Parallelism✅ Yes (via native C++ WorkerPool)✅ Yes (via native Rust Rayon multi-threading)⚠️ Partial (via Python ThreadPoolExecutor mapping to Rust)Within-Doc Parallelism✅ Yes (parallel_encode)❌ No (sequential per document)❌ No (sequential per document)Training from Iterator✅ Yes (sentence_iterator)✅ Yes (train_from_iterator)❌ NoSubword Regularization (Sampling / N-best)✅ Yes (at call-time; supports Unigram sampling and BPE-dropout)⚠️ Partial (BPE-dropout only, configured at model training/load time, no call-time sampling)❌ NoIn-Memory Add Tokens⚠️ Partial (Intentional design: dynamic adding is not supported; requires model recreation via modified protobuf)✅ Yes (Easy, via add_tokens)⚠️ Partial (Intentional design: dynamic adding not supported; requires model recreation via modified dicts)Pre-Tokenization (Modular)❌ No (Intentional design: parses raw Unicode stream without pre-splitting)✅ Yes (fully modular: tokenizer.pre_tokenizer = ...)⚠️ Partial (static via regex pat_str in constructor)Modular Post-Processing⚠️ Partial (BOS/EOS toggles and extra options only)✅ Yes (template post-processor)❌ NoText Normalization Support✅ Yes (baked into model, highly optimized precompiled rules)✅ Yes (fully modular pipeline, e.g. Lowercase, NFKC)❌ No (Intentional design: tokenizes raw input exactly as-is)Custom Normalizers✅ Yes (defined at training time via TSV mapping, or at runtime via Python mapping passed to trainer)✅ Yes (custom Python/Rust functions or sequences)❌ NoSpecial Token Security Policy✅ Yes (static per-token: control_symbols [safe] vs user_defined_symbols [unsafe] in same model)⚠️ Partial (global toggle split_special_tokens, cannot mix per-token)⚠️ Partial (global call-time toggle allowed_special/disallowed_special, cannot mix per-token)Token/Char Alignment Helpers❌ No (returns raw offsets only)✅ Yes (char_to_token(), token_to_word(), etc.)❌ No
———
2. Usage & Code Comparison
Below is a side-by-side comparison of common tasks across the three libraries.
2.1. Instantiate
SentencePiece: importsentencepieceasspm# Load from filesp=spm.SentencePieceProcessor.from_file("m.model") # Load from in-memory bytessp=spm.SentencePieceProcessor.from_proto(proto_bytes)
Hugging Face tokenizers: fromtokenizersimportTokenizertokenizer=Tokenizer.from_file("t.json")
tiktoken: importtiktokenenc=tiktoken.get_encoding("cl100k_base") # or:enc=tiktoken.encoding_for_model("gpt-4")
2.2. Encode (IDs)
SentencePiece: ids=sp.encode("Hello world", return_type=int)
Hugging Face tokenizers: ids=tokenizer.encode("Hello world").ids
tiktoken: ids=enc.encode("Hello world")
2.3. Encode (Pieces)
SentencePiece: pieces=sp.encode("Hello world", return_type=str) # Returns list[str]: ['▁Hello', '▁world']# Or to get raw bytes pieces directly:pieces_bytes=sp.encode("Hello world", return_type=bytes) # Returns list[bytes]: [b'\xe2\x96\x81Hello', b'\xe2\x96\x81world']
Hugging Face tokenizers: pieces=tokenizer.encode("Hello world").tokens# Returns list[str]
tiktoken: # tiktoken does not support direct token-to-string piece representation# without decoding each ID individually to bytes:pieces= [enc.decode_single_token_bytes(t) fortinids] # Returns list[bytes]
2.4. Decode (Str)
SentencePiece: text=sp.decode(ids)
Hugging Face tokenizers: text=tokenizer.decode(ids)
tiktoken: text=enc.decode(ids)
2.5. Decode (Bytes)
SentencePiece: text_bytes=sp.decode(ids, return_type=bytes)
Hugging Face tokenizers: Not Supported (must manually decode string to bytes)
tiktoken: text_bytes=enc.decode_bytes(ids)
2.6. Offset Mapping (Unicode)
SentencePiece: # Encode with Unicode character offsets:res=sp.encode("text", return_type='offset_mapping') # Or force Unicode offsets on bytes input explicitly:res=sp.encode(b"text", return_type='offset_mapping', return_bytes=False) # res['offsets'] -> list[tuple[int, int]]# Decode with Unicode character offsets:res=sp.decode(ids, return_type='offset_mapping') # res['text'] -> str, res['offsets'] -> list[tuple[int, int]]
Hugging Face tokenizers: # Encode with Unicode character offsets:offsets=tokenizer.encode("text").offsets# offsets -> list[tuple[int, int]]# Decode with offsets:# *Not Supported*
tiktoken: Encode: Not Supported
# Decode with offsets:text, offsets=enc.decode_with_offsets(ids) # text -> str, offsets -> list[int] (start character indices)
2.7. Offset Mapping (Bytes)
SentencePiece: # Encode with raw byte offsets (triggered by bytes input):res=sp.encode(b"text", return_type='offset_mapping') # Or force raw byte offsets on str input explicitly:res=sp.encode("text", return_type='offset_mapping', return_bytes=True) # res['offsets'] -> list[tuple[int, int]] (byte offsets)# Decode with raw byte offsets (explicit return_bytes=True):res=sp.decode(ids, return_type='offset_mapping', return_bytes=True) # res['text'] -> bytes, res['offsets'] -> list[tuple[int, int]] (byte offsets)
Hugging Face tokenizers: Not Supported
tiktoken: Not Supported
2.8. NumPy Support
SentencePiece: # Direct encoding to NumPy array (zero-copy):arr=sp.encode("text", return_type='numpy') # Decoding directly from NumPy array:text=sp.decode(arr)
Hugging Face tokenizers: importnumpyasnp# Requires manual conversion:arr=np.array(tokenizer.encode("text").ids)
tiktoken: # Direct encoding to NumPy array (zero-copy):arr=enc.encode_to_numpy("text")
2.9. Encode Batch
SentencePiece: # Parallel processing in C++ (releasing GIL):ids_list=sp.encode(texts, num_threads=4)
Hugging Face tokenizers: # Parallel processing in Rust (Rayon):outputs=tokenizer.encode_batch(texts) ids_list= [o.idsforoinoutputs]
tiktoken: # Parallel processing via Python ThreadPoolExecutor mapping to Rust:ids_list=enc.encode_batch(texts, num_threads=4)
2.10. Parallel (Within-Doc)
SentencePiece: # Parallelize encoding of a single large document:pool=spm.ThreadPool(num_threads=4) ids=sp.parallel_encode(large_text, chunk_len=1048576, thread_pool=pool, return_type=int)
Hugging Face tokenizers: Not Supported (encodes single document sequentially on one thread)
tiktoken: Not Supported (encodes single document sequentially on one thread)
2.11. Thread Pool
SentencePiece: # Explicit thread pool sharing across multiple batch calls:pool=spm.ThreadPool(num_threads=8) ids=sp.encode(batch, thread_pool=pool)
Hugging Face tokenizers: Managed internally in Rust via global Rayon thread pool.
tiktoken: Managed internally in Rust thread pool.
2.12. Train from Iterator
SentencePiece: spm.SentencePieceTrainer.train(sentence_iterator=my_iter, model_prefix='m', vocab_size=1000)
Hugging Face tokenizers: fromtokenizers.trainersimportBpeTrainertokenizer.train_from_iterator(my_iter, trainer=BpeTrainer())
tiktoken: Not Supported (training is not exposed in public Python API)
2.13. Model Modification (In-Memory)
SentencePiece: # Intentional design: The processor is immutable.# To modify, rewrite the model protobuf and recreate the processor:importsentencepiece_model_pb2aspbproto=pb.ModelProto() proto.ParseFromString(open("m.model", "rb").read()) # ... modify proto (e.g. proto.pieces.add()) ...sp=spm.SentencePieceProcessor.from_proto(proto.SerializeToString())
Hugging Face tokenizers: # Dynamically add tokens to loaded model:tokenizer.add_tokens(["<|my_token|>"]) tokenizer.add_special_tokens(["<|special|>"])
tiktoken: # Recreate encoding with modified ranks/special tokens dicts:enc=tiktoken.get_encoding("cl100k_base") ranks=dict(enc._mergeable_ranks) special=dict(enc._special_tokens) # Modify dictsranks[b"new_token"] =len(ranks) new_enc=tiktoken.Encoding( name="modified_cl100k", pat_str=enc._pat_str, mergeable_ranks=ranks, special_tokens=special )
2.14. Configure Pre-tokenizer
SentencePiece: Not Supported dynamically (pre-tokenization is baked into model normalization at training time)
Hugging Face tokenizers: fromtokenizersimportpre_tokenizers# Configure modular pre-tokenizer:tokenizer.pre_tokenizer=pre_tokenizers.Whitespace()
tiktoken: Partial (static regex pattern passed during custom class initialization)
2.15. Special Token Security
SentencePiece: # Defined statically at training time:spm.train(..., control_symbols=['<c>'], user_defined_symbols=['<u>']) # <c> is safe (never split from raw text), <u> is parsed from text.
Hugging Face tokenizers: # Configure behavior at load/call time:tokenizer=AutoTokenizer.from_pretrained(..., split_special_tokens=True)
tiktoken: # Enforced at call time:enc.encode("text <|endoftext|>", allowed_special=set(), disallowed_special="all") # Raises ValueError on unexpected special token injection
2.16. Alignment Helpers
SentencePiece: Not Supported (returns raw offsets, but no high-level helper APIs to map char to token)
Hugging Face tokenizers: output=tokenizer.encode("text") token_id=output.char_to_token(5) span=output.token_to_chars(token_id)
tiktoken: Not Supported
2.17. Configure Normalizer
SentencePiece: Training-time only (normalization rule defined via normalization_rule_name or normalization_rule_tsv during training). You can also pass a pre-configured SentencePieceNormalizer instance: # Define normalizer rules at runtime in Pythonnorm=spm.SentencePieceNormalizer(norm_map=[('foo', 'bar')], escape_whitespaces=True) # Pass the normalizer instance to the trainerspm.SentencePieceTrainer.train(..., normalizer=norm)
Hugging Face tokenizers: fromtokenizersimportnormalizers# Configure normalizer pipeline dynamically:tokenizer.normalizer=normalizers.Sequence([normalizers.NFKC(), normalizers.Lowercase()])
tiktoken: Not Supported (tokenizes raw input exactly as-is)
2.18. Independent Normalization
SentencePiece: # Normalize text independently of tokenization:norm=spm.SentencePieceNormalizer(model_file='m.model') print(norm.normalize("text")) # Or via processor:print(sp.normalize("text")) # Or initialize with custom mapping directly:norm_custom=spm.SentencePieceNormalizer(norm_map=[('foo', 'bar')]) print(norm_custom.normalize("foo")) # Output: bar
Hugging Face tokenizers: # Normalize text independently:print(tokenizer.normalizer.normalize_str("text"))
tiktoken: Not Supported
2.19. Normalization Offsets
SentencePiece: # Retrieve character-to-byte offset mapping after normalization:normalized, offsets=sp.normalize("text", with_offsets=True)
Hugging Face tokenizers: Not Supported directly on normalizer (offsets are computed during full encoding only)
tiktoken: Not Supported
2.20. Subword Regularization / Sampling (N-best)
SentencePiece: # Encode with sampling (Unigram sampling or BPE-dropout depending on model)# nbest_size: -1 (unlimited, default), alpha: 0.1 (smoothing parameter)ids=sp.encode("Hello world", enable_sampling=True, alpha=0.1, nbest_size=-1) # N-best encoding (returns list of list of segmentations)nbest_ids=sp.nbest_encode("Hello world", nbest_size=5, return_type=int) # nbest_ids -> list[list[int]] (5 segmentations)
Hugging Face tokenizers: # BPE-dropout is configured at model creation time, not at encode call time:fromtokenizers.modelsimportBPEmodel=BPE(dropout=0.1) # 10% dropout during tokenization# Unigram sampling / N-best: Not Supported in fast tokenizers API
tiktoken: Not Supported
2.21. Decode NumPy Input
SentencePiece: importnumpyasnp# Decode 1D NumPy array (single sequence)arr_1d=np.array([284, 47, 11], dtype=np.int32) text=sp.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[284, 47, 11], [10, 20, 30]], dtype=np.int32) texts=sp.decode(arr_2d)
Hugging Face tokenizers: importnumpyasnp# Decode 1D NumPy arrayarr_1d=np.array([101, 7592, 102], dtype=np.uint32) text=tokenizer.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[101, 7592, 102], [101, 2088, 102]], dtype=np.uint32) texts=tokenizer.decode_batch(arr_2d)
tiktoken: importnumpyasnp# Decode 1D NumPy arrayarr_1d=np.array([31373, 995], dtype=np.uint32) text=enc.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[31373, 995], [11274, 16390]], dtype=np.uint32) texts=enc.decode_batch(arr_2d)