Prompt Cache Hit Rate Calculator

Prompt Cache Hit Rate Calculator MCP Connector for Claude

A+

Analyze LLM prompt prefix caching efficiency and performance.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides deterministic diagnostic tools to analyze the effectiveness of LLM prompt prefix caching. Use analyze_cache_performance to calculate hit rates, tokens saved, and cache efficiency. You can also use calculate_warmup_metrics to determine how long it takes to reach cache saturation, or validate_cache_configuration to check if your current TTL and cache size settings are optimal for your request patterns.

llmcachingperformancediagnosticstokens

3 tools expose this connector's capabilities to your AI agent.

calculate_warmup_metrics

validate_cache_configuration

analyze_cache_performance

See how to talk to your AI agent using Prompt Cache Hit Rate Calculator.

Analyze these request logs with a TTL of 300 seconds and a cache size of 1000 tokens.

The analysis shows a cache hit rate of 0.45, with 450 tokens saved and a cache efficiency of 0.12.

How long does it take to fill a 5000 token cache based on these logs?

The cache reaches saturation in 1240 seconds.

Are my current cache settings optimal for this workload?

No, the current settings are not optimal. It is suggested to increase the cache size to 2500 tokens to improve the hit rate.

You can use the `analyze_cache_performance` tool. It processes your request logs and returns the exact `cacheHits` and `cacheHitRate` based on your provided TTL and cache size.

Related Connectors