Tokenization Normalizer

Tokenization Normalizer MCP Connector for Claude

A+

Resolves tokenization drift by normalizing text to match specific LLM tokenizer profiles.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides tools to resolve tokenization drift, ensuring deterministic control planes can accurately predict LLM behavior. By applying strict transformation rules for specific profiles like cl100k_base or llama, it adjusts whitespace, merges contractions, and normalizes Unicode. Use normalize_text_structure to transform text, analyze_tokenization_drift to assess instability, or get_tokenizer_capabilities to inspect supported rules.

tokenizationllmnormalizationdrift-detectiondeterministic

3 tools expose this connector's capabilities to your AI agent.

get_tokenizer_capabilities

Informs the user about what specific normalization rules are available for a given tokenizer profile

analyze_tokenization_drift

Compares the original text against a target profile to provide a summary of how "unstable" the text is for that specific model

normalize_text_structure

Transforms a given string to match the structural expectations of a specific target tokenizer

See how to talk to your AI agent using Tokenization Normalizer.

How much drift is in this text for the llama tokenizer: 'Hello, world!'?

The drift severity for 'Hello, world!' using the llama profile is low, with a primary transformation needed in whitespace alignment.

Normalize this text for cl100k_base: 'Don't stop now.'

The normalized text is 'Don't stop now.' and the estimated token count delta is 0.

What rules does the llama tokenizer support?

The llama tokenizer supports unicode_standardization and contraction_merging rules.

Tokenization drift occurs when the same text results in different token counts or boundaries across different LLM models. This tool helps align text to a target profile to minimize these differences.

Related Connectors