Multi-Language Token Estimator

Multi-Language Token Estimator MCP Connector for Claude

A+

Analyze text composition and estimate token counts across multiple languages.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides deterministic linguistic analysis for mixed-language text. It uses character-level Unicode classification to identify English, Chinese, Japanese, Korean, Cyrillic, and Arabic content. Use estimate_text_composition to get a detailed breakdown of character counts, token estimates, and linguistic dominance, or get_language_ratios to view the fixed tokenization multipliers used for each language.

tokensunicodemultilingualtext-analysisllm-optimization

3 tools expose this connector's capabilities to your AI agent.

estimate_text_composition

Analyzes a text string to provide a granular breakdown of character counts, token estimates, and linguistic dominance

validate_unicode_range

Verifies if a specific character belongs to one of the supported linguistic ranges

get_language_ratios

Retrieves the current deterministic token-to-character ratios used by the system

See how to talk to your AI agent using Multi-Language Token Estimator.

Analyze the composition of this text: 'Hello 世界'

{ "language_breakdown": [ { "language": "english", "chars": 5, "tokens": 20.0 }, { "language": "chinese", "chars": 2, "tokens": 3.0 } ], "total_tokens": 23.0, "dominant_language": "english", "language_mix_ratio": 0.714 }

What are the tokenization ratios used by this tool?

{ "ratios": { "english": 4.0, "chinese": 1.5, "japanese": 1.4, "korean": 1.8, "cyrillic": 2.8, "arabic": 2.5 } }

Is the character 'あ' supported?

{ "language": "japanese", "isValid": true }

Tokens are estimated by multiplying the character count of each language by a specific deterministic ratio (e.g., 4.0 for English, 1.5 for Chinese).

Related Connectors