WordPieceEmbeddingTokenizer class

EmbeddingTokenizer for BERT-style WordPiece models. encode(prefix, text) tokenizes prefix + text (matching the SentencePiece adapter's convention) as [CLS] + wordpiece(prefix+text) + [SEP], with an all-ones TokenizedInput.attentionMask and all-zeros TokenizedInput.tokenTypeIds (single-segment input).

Implemented types

Properties

clsId int
[CLS] special-token id (BERT convention: 101).
final
hashCode int
The hash code for this object.
no setterinherited
runtimeType Type
A representation of the runtime type of the object.
no setterinherited
sepId int
[SEP] special-token id (BERT convention: 102).
final
unkId int
[UNK] special-token id (BERT convention: 100).
final

Methods

encode(String prefix, String text) TokenizedInput
Tokenizes (prefix + text) per this tokenizer's convention (BOS/EOS wrap for SentencePiece, CLS/.../SEP for WordPiece).
override
noSuchMethod(Invocation invocation) → dynamic
Invoked when a nonexistent method or property is accessed.
inherited
toString() String
A string representation of this object.
inherited

Operators

operator ==(Object other) bool
The equality operator.
inherited

Static Methods

fromJsonString(String content) WordPieceEmbeddingTokenizer
Parse a WordPiece tokenizer from the raw tokenizer.json content.
fromPath(String path) Future<WordPieceEmbeddingTokenizer>
Load a WordPiece tokenizer from a HuggingFace tokenizer.json at path.
isWordPieceJson(Map<String, dynamic> json) bool
Returns true when json (a decoded tokenizer.json map) describes a WordPiece model. Used by the ONNX backend's tokenizer loader to route between SentencePiece and WordPiece without a file-extension heuristic.