KV Cache Optimizer

KV Cache Optimizer MCP Connector for Claude

A+

Deterministic calculator for LLM KV cache memory, hardware utilization, and performance impact.

4 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides precise tools for estimating Large Language Model (LLM) KV cache memory consumption and hardware requirements. It allows users to calculate the exact memory footprint using calculate_kv_cache_footprint, evaluate memory savings with analyze_optimization_strategy (supporting sliding window and paged attention), and verify hardware compatibility via evaluate_hardware_feasibility. Additionally, users can determine the maximum efficient workload using optimize_batch_configuration to maximize throughput within GPU memory constraints.

kv-cachellm-inferencegpu-memorypaged-attentionthroughput

4 tools expose this connector's capabilities to your AI agent.

analyze_optimization_strategy

Evaluates how much memory is saved or how efficiently memory is used when applying specific optimization techniques

calculate_kv_cache_footprint

Determines the total memory required to store the KV cache for a specific model configuration and batch

evaluate_hardware_feasibility

Checks if the requested workload fits within the physical constraints of the target GPU

optimize_batch_configuration

Finds the highest possible batch size that does not violate the memory constraints

See how to talk to your AI agent using KV Cache Optimizer.

Calculate the KV cache footprint for a model with 32 layers, 32 heads, 128 head dimension, 2048 sequence length, and a batch size of 8 using FP16.

The total KV cache footprint for this configuration is 128 GiB.

What is the optimal batch size for a model with 32 layers, 32 heads, 128 head dimension, 2048 sequence length, 40GB of GPU memory, FP16 precision, and 0.5 seconds latency?

The optimal batch size is 2, providing a throughput impact of 4.0 tokens per second.

Check if a 100GB KV cache fits in a GPU with 80GB of VRAM.

The workload is not feasible as the requested cache size exceeds the available GPU memory.

It provides deterministic calculations for KV cache size, memory bandwidth requirements, and optimal batch sizes, helping you avoid Out-of-Memory errors.

Related Connectors