sentencepiece/python/tokenizer_comparison_cheat_sheet.md at master · google/sentencepiece

GitHub

Tokenizer Comparison Cheat Sheet

This cheat sheet compares the capabilities, API designs, and performance features of three major tokenizer libraries: SentencePiece (new pybind11 API), Hugging Face Tokenizers, and OpenAI's tiktoken.

Note: Please feel free to send a PR if you find any problems or inaccuracies.

———
1. Capability Matrix

Feature / CapabilitySentencePieceHugging Face tokenizerstiktokenCompared Version>=v0.2.2v0.23.1v0.13.0Native BackendC++ / pybind11Rust / PyO3Rust / PyO3Supported AlgorithmsBPE, Unigram, Char, WordBPE, WordPiece, Unigram, WordLevelBPEOOV / Unknown Handling<unk> token, byte_fallback<unk> token, byte_fallback (BPE/Unigram), or Byte-level BPE (no <unk>)Byte-level BPE (no <unk> token)Training Support✅ Yes (via SentencePieceTrainer)✅ Yes (via Trainer classes)❌ NoGIL Release (Parallelism)✅ Yes (Releases GIL in almost all cases including single/batch encode & decode; supports Free-Threading)⚠️ Partial (Releases GIL during core Rust execution, but string copying is done with GIL held)⚠️ Partial (Releases GIL during core Rust execution, but python-level threading and conversions hold GIL)Encode Offset Mapping (Unicode)✅ Yes (via return_type='offset_mapping')✅ Yes (via Encoding.offsets)❌ NoDecode Offset Mapping (Unicode)✅ Yes (via return_type='offset_mapping', includes text key)❌ No⚠️ Partial (via Python decode_with_offsets; start offsets only)Raw Byte Offset Mapping✅ Yes (triggered by bytes input or return_bytes=True for both encode and decode)❌ No❌ NoDirect UTF-8 Bytes I/O✅ Yes (accepts bytes for encode, returns bytes for decode)❌ No (requires str)⚠️ Partial (requires str for encode, returns bytes via decode_bytes)Zero-Copy Input (str/bytes)✅ Yes (zero-copy for both str and bytes via Pybind11 memory view)❌ No (always copies input str to Rust owned String via PyO3)⚠️ Partial (zero-copy for str under some conditions, no bytes support)NumPy Array Output (Encode)✅ Yes (Direct to NumPy, zero-copy)❌ No (Transformers wrapper only)✅ Yes (via encode_to_numpy, zero-copy)NumPy Array Input (Decode)✅ Yes (Accepts 1D/2D; zero-copy on best-effort basis)✅ Yes (Accepts 1D/2D; copied)✅ Yes (Accepts 1D/2D; copied)Batch Encoding Parallelism✅ Yes (via native C++ WorkerPool)✅ Yes (via native Rust Rayon multi-threading)⚠️ Partial (via Python ThreadPoolExecutor mapping to Rust)Within-Doc Parallelism✅ Yes (parallel_encode)❌ No (sequential per document)❌ No (sequential per document)Training from Iterator✅ Yes (sentence_iterator)✅ Yes (train_from_iterator)❌ NoSubword Regularization (Sampling / N-best)✅ Yes (at call-time; supports Unigram sampling and BPE-dropout)⚠️ Partial (BPE-dropout only, configured at model training/load time, no call-time sampling)❌ NoIn-Memory Add Tokens⚠️ Partial (Intentional design: dynamic adding is not supported; requires model recreation via modified protobuf)✅ Yes (Easy, via add_tokens)⚠️ Partial (Intentional design: dynamic adding not supported; requires model recreation via modified dicts)Pre-Tokenization (Modular)❌ No (Intentional design: parses raw Unicode stream without pre-splitting)✅ Yes (fully modular: tokenizer.pre_tokenizer = ...)⚠️ Partial (static via regex pat_str in constructor)Modular Post-Processing⚠️ Partial (BOS/EOS toggles and extra options only)✅ Yes (template post-processor)❌ NoText Normalization Support✅ Yes (baked into model, highly optimized precompiled rules)✅ Yes (fully modular pipeline, e.g. Lowercase, NFKC)❌ No (Intentional design: tokenizes raw input exactly as-is)Custom Normalizers✅ Yes (defined at training time via TSV mapping, or at runtime via Python mapping passed to trainer)✅ Yes (custom Python/Rust functions or sequences)❌ NoSpecial Token Security Policy✅ Yes (static per-token: control_symbols [safe] vs user_defined_symbols [unsafe] in same model)⚠️ Partial (global toggle split_special_tokens, cannot mix per-token)⚠️ Partial (global call-time toggle allowed_special/disallowed_special, cannot mix per-token)Token/Char Alignment Helpers❌ No (returns raw offsets only)✅ Yes (char_to_token(), token_to_word(), etc.)❌ No
———
2. Usage & Code Comparison

Below is a side-by-side comparison of common tasks across the three libraries.

2.1. Instantiate

SentencePiece: importsentencepieceasspm# Load from filesp=spm.SentencePieceProcessor.from_file("m.model") # Load from in-memory bytessp=spm.SentencePieceProcessor.from_proto(proto_bytes)

Hugging Face tokenizers: fromtokenizersimportTokenizertokenizer=Tokenizer.from_file("t.json")

tiktoken: importtiktokenenc=tiktoken.get_encoding("cl100k_base") # or:enc=tiktoken.encoding_for_model("gpt-4")

2.2. Encode (IDs)

SentencePiece: ids=sp.encode("Hello world", return_type=int)

Hugging Face tokenizers: ids=tokenizer.encode("Hello world").ids

tiktoken: ids=enc.encode("Hello world")

2.3. Encode (Pieces)

SentencePiece: pieces=sp.encode("Hello world", return_type=str) # Returns list[str]: ['▁Hello', '▁world']# Or to get raw bytes pieces directly:pieces_bytes=sp.encode("Hello world", return_type=bytes) # Returns list[bytes]: [b'\xe2\x96\x81Hello', b'\xe2\x96\x81world']

Hugging Face tokenizers: pieces=tokenizer.encode("Hello world").tokens# Returns list[str]

tiktoken: # tiktoken does not support direct token-to-string piece representation# without decoding each ID individually to bytes:pieces= [enc.decode_single_token_bytes(t) fortinids] # Returns list[bytes]

2.4. Decode (Str)

SentencePiece: text=sp.decode(ids)

Hugging Face tokenizers: text=tokenizer.decode(ids)

tiktoken: text=enc.decode(ids)

2.5. Decode (Bytes)

SentencePiece: text_bytes=sp.decode(ids, return_type=bytes)

Hugging Face tokenizers: Not Supported (must manually decode string to bytes)

tiktoken: text_bytes=enc.decode_bytes(ids)

2.6. Offset Mapping (Unicode)

SentencePiece: # Encode with Unicode character offsets:res=sp.encode("text", return_type='offset_mapping') # Or force Unicode offsets on bytes input explicitly:res=sp.encode(b"text", return_type='offset_mapping', return_bytes=False) # res['offsets'] -> list[tuple[int, int]]# Decode with Unicode character offsets:res=sp.decode(ids, return_type='offset_mapping') # res['text'] -> str, res['offsets'] -> list[tuple[int, int]]

Hugging Face tokenizers: # Encode with Unicode character offsets:offsets=tokenizer.encode("text").offsets# offsets -> list[tuple[int, int]]# Decode with offsets:# *Not Supported*

tiktoken: Encode: Not Supported

# Decode with offsets:text, offsets=enc.decode_with_offsets(ids) # text -> str, offsets -> list[int] (start character indices)

2.7. Offset Mapping (Bytes)

SentencePiece: # Encode with raw byte offsets (triggered by bytes input):res=sp.encode(b"text", return_type='offset_mapping') # Or force raw byte offsets on str input explicitly:res=sp.encode("text", return_type='offset_mapping', return_bytes=True) # res['offsets'] -> list[tuple[int, int]] (byte offsets)# Decode with raw byte offsets (explicit return_bytes=True):res=sp.decode(ids, return_type='offset_mapping', return_bytes=True) # res['text'] -> bytes, res['offsets'] -> list[tuple[int, int]] (byte offsets)

Hugging Face tokenizers: Not Supported

tiktoken: Not Supported

2.8. NumPy Support

SentencePiece: # Direct encoding to NumPy array (zero-copy):arr=sp.encode("text", return_type='numpy') # Decoding directly from NumPy array:text=sp.decode(arr)

Hugging Face tokenizers: importnumpyasnp# Requires manual conversion:arr=np.array(tokenizer.encode("text").ids)

tiktoken: # Direct encoding to NumPy array (zero-copy):arr=enc.encode_to_numpy("text")

2.9. Encode Batch

SentencePiece: # Parallel processing in C++ (releasing GIL):ids_list=sp.encode(texts, num_threads=4)

Hugging Face tokenizers: # Parallel processing in Rust (Rayon):outputs=tokenizer.encode_batch(texts) ids_list= [o.idsforoinoutputs]

tiktoken: # Parallel processing via Python ThreadPoolExecutor mapping to Rust:ids_list=enc.encode_batch(texts, num_threads=4)

2.10. Parallel (Within-Doc)

SentencePiece: # Parallelize encoding of a single large document:pool=spm.ThreadPool(num_threads=4) ids=sp.parallel_encode(large_text, chunk_len=1048576, thread_pool=pool, return_type=int)

Hugging Face tokenizers: Not Supported (encodes single document sequentially on one thread)

tiktoken: Not Supported (encodes single document sequentially on one thread)

2.11. Thread Pool

SentencePiece: # Explicit thread pool sharing across multiple batch calls:pool=spm.ThreadPool(num_threads=8) ids=sp.encode(batch, thread_pool=pool)

Hugging Face tokenizers: Managed internally in Rust via global Rayon thread pool.

tiktoken: Managed internally in Rust thread pool.

2.12. Train from Iterator

SentencePiece: spm.SentencePieceTrainer.train(sentence_iterator=my_iter, model_prefix='m', vocab_size=1000)

Hugging Face tokenizers: fromtokenizers.trainersimportBpeTrainertokenizer.train_from_iterator(my_iter, trainer=BpeTrainer())

tiktoken: Not Supported (training is not exposed in public Python API)

2.13. Model Modification (In-Memory)

SentencePiece: # Intentional design: The processor is immutable.# To modify, rewrite the model protobuf and recreate the processor:importsentencepiece_model_pb2aspbproto=pb.ModelProto() proto.ParseFromString(open("m.model", "rb").read()) # ... modify proto (e.g. proto.pieces.add()) ...sp=spm.SentencePieceProcessor.from_proto(proto.SerializeToString())

Hugging Face tokenizers: # Dynamically add tokens to loaded model:tokenizer.add_tokens(["<|my_token|>"]) tokenizer.add_special_tokens(["<|special|>"])

tiktoken: # Recreate encoding with modified ranks/special tokens dicts:enc=tiktoken.get_encoding("cl100k_base") ranks=dict(enc._mergeable_ranks) special=dict(enc._special_tokens) # Modify dictsranks[b"new_token"] =len(ranks) new_enc=tiktoken.Encoding( name="modified_cl100k", pat_str=enc._pat_str, mergeable_ranks=ranks, special_tokens=special )

2.14. Configure Pre-tokenizer

SentencePiece: Not Supported dynamically (pre-tokenization is baked into model normalization at training time)

Hugging Face tokenizers: fromtokenizersimportpre_tokenizers# Configure modular pre-tokenizer:tokenizer.pre_tokenizer=pre_tokenizers.Whitespace()

tiktoken: Partial (static regex pattern passed during custom class initialization)

2.15. Special Token Security

SentencePiece: # Defined statically at training time:spm.train(..., control_symbols=['<c>'], user_defined_symbols=['<u>']) # <c> is safe (never split from raw text), <u> is parsed from text.

Hugging Face tokenizers: # Configure behavior at load/call time:tokenizer=AutoTokenizer.from_pretrained(..., split_special_tokens=True)

tiktoken: # Enforced at call time:enc.encode("text <|endoftext|>", allowed_special=set(), disallowed_special="all") # Raises ValueError on unexpected special token injection

2.16. Alignment Helpers

SentencePiece: Not Supported (returns raw offsets, but no high-level helper APIs to map char to token)

Hugging Face tokenizers: output=tokenizer.encode("text") token_id=output.char_to_token(5) span=output.token_to_chars(token_id)

tiktoken: Not Supported

2.17. Configure Normalizer

SentencePiece: Training-time only (normalization rule defined via normalization_rule_name or normalization_rule_tsv during training). You can also pass a pre-configured SentencePieceNormalizer instance: # Define normalizer rules at runtime in Pythonnorm=spm.SentencePieceNormalizer(norm_map=[('foo', 'bar')], escape_whitespaces=True) # Pass the normalizer instance to the trainerspm.SentencePieceTrainer.train(..., normalizer=norm)

Hugging Face tokenizers: fromtokenizersimportnormalizers# Configure normalizer pipeline dynamically:tokenizer.normalizer=normalizers.Sequence([normalizers.NFKC(), normalizers.Lowercase()])

tiktoken: Not Supported (tokenizes raw input exactly as-is)

2.18. Independent Normalization

SentencePiece: # Normalize text independently of tokenization:norm=spm.SentencePieceNormalizer(model_file='m.model') print(norm.normalize("text")) # Or via processor:print(sp.normalize("text")) # Or initialize with custom mapping directly:norm_custom=spm.SentencePieceNormalizer(norm_map=[('foo', 'bar')]) print(norm_custom.normalize("foo")) # Output: bar

Hugging Face tokenizers: # Normalize text independently:print(tokenizer.normalizer.normalize_str("text"))

tiktoken: Not Supported

2.19. Normalization Offsets

SentencePiece: # Retrieve character-to-byte offset mapping after normalization:normalized, offsets=sp.normalize("text", with_offsets=True)

Hugging Face tokenizers: Not Supported directly on normalizer (offsets are computed during full encoding only)

tiktoken: Not Supported

2.20. Subword Regularization / Sampling (N-best)

SentencePiece: # Encode with sampling (Unigram sampling or BPE-dropout depending on model)# nbest_size: -1 (unlimited, default), alpha: 0.1 (smoothing parameter)ids=sp.encode("Hello world", enable_sampling=True, alpha=0.1, nbest_size=-1) # N-best encoding (returns list of list of segmentations)nbest_ids=sp.nbest_encode("Hello world", nbest_size=5, return_type=int) # nbest_ids -> list[list[int]] (5 segmentations)

Hugging Face tokenizers: # BPE-dropout is configured at model creation time, not at encode call time:fromtokenizers.modelsimportBPEmodel=BPE(dropout=0.1) # 10% dropout during tokenization# Unigram sampling / N-best: Not Supported in fast tokenizers API

tiktoken: Not Supported

2.21. Decode NumPy Input

SentencePiece: importnumpyasnp# Decode 1D NumPy array (single sequence)arr_1d=np.array([284, 47, 11], dtype=np.int32) text=sp.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[284, 47, 11], [10, 20, 30]], dtype=np.int32) texts=sp.decode(arr_2d)

Hugging Face tokenizers: importnumpyasnp# Decode 1D NumPy arrayarr_1d=np.array([101, 7592, 102], dtype=np.uint32) text=tokenizer.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[101, 7592, 102], [101, 2088, 102]], dtype=np.uint32) texts=tokenizer.decode_batch(arr_2d)

tiktoken: importnumpyasnp# Decode 1D NumPy arrayarr_1d=np.array([31373, 995], dtype=np.uint32) text=enc.decode(arr_1d) # Decode 2D NumPy array (batch)arr_2d=np.array([[31373, 995], [11274, 16390]], dtype=np.uint32) texts=enc.decode_batch(arr_2d)