Prompt Cache Hit Calculator

Prompt Cache Hit Calculator MCP Connector for Claude

A+

Analyze prompt prefix caching performance, efficiency, and cost savings.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides deterministic analysis of prompt prefix caching strategies. It allows AI agents to evaluate how effectively cached prefixes are being utilized to reduce latency and costs. Use analyze_cache_performance to calculate hit rates and monetary savings, evaluate_cache_optimization to determine the ideal cache size based on prefix sharing, and inspect_cache_dynamics to monitor eviction rates and specific prefix match lengths.

cachingtokenscost-analysisprompt-engineeringperformance

3 tools expose this connector's capabilities to your AI agent.

analyze_cache_performance

Provides a high-level overview of how well the cache is performing regarding hits, efficiency, and cost savings

inspect_cache_dynamics

Investigates the frequency of cache turnover and the specific overlap between individual requests

evaluate_cache_optimization

Identifies the ideal cache capacity and the degree of prefix overlap to guide infrastructure scaling

See how to talk to your AI agent using Prompt Cache Hit Calculator.

Analyze my cache performance with a cache size of 1000 tokens and a cost of 0.00002 per token.

The cache hit rate is 0.45, with a cache efficiency of 0.25. The total cache hit value saved is $0.12.

What is the optimal cache size for these request logs?

The optimal cache size is 1540 tokens, with a prefix sharing ratio of 0.65.

Check the cache eviction rate for a 5000 token cache.

The cache eviction rate is 0.02, indicating a stable cache with low turnover.

You can use the `analyze_cache_performance` tool. By providing the `costPerToken` parameter, the tool calculates the total `cacheHitValue` based on the tokens saved during successful hits.

Related Connectors