Agent Evaluation Metrics Calculator

Agent Evaluation Metrics Calculator MCP Connector for Claude

A+

Quantify agent performance with deterministic accuracy, speed, and efficiency metrics.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides a deterministic engine to evaluate AI agent performance. It calculates core success metrics like accuracy, precision, recall, and F1 score using calculate_core_metrics. It also analyzes operational costs through calculate_efficiency_metrics to determine latency and token efficiency. Finally, it generates a weighted composite score and detects performance regressions via calculate_composite_and_health based on a defined baseline accuracy.

metricsaccuracylatencyefficiencyevaluation

3 tools expose this connector's capabilities to your AI agent.

calculate_composite_and_health

Generates a single performance score and determines if the agent is underperforming relative to a baseline

calculate_efficiency_metrics

Analyzes the operational cost and speed of the agent

calculate_core_metrics

Calculates the foundational performance indicators (accuracy, precision, recall, and F1 score) based on task outcomes

See how to talk to your AI agent using Agent Evaluation Metrics Calculator.

Calculate the core metrics for these tasks: [{'task_id': '1', 'expected_output': 'A', 'actual_output': 'A', 'latency_ms': 100, 'tokens_used': 50, 'success_boolean': true}, {'task_id': '2', 'expected_output': 'B', 'actual_output': 'C', 'latency_ms': 150, 'tokens_used': 60, 'success_boolean': false}]

{"accuracy": 0.5, "precision": 1.0, "recall": 0.5, "f1Score": 0.6666666666666666}

What is the token efficiency for 10 successful tasks using 500 tokens?

0.02

Check if an agent with 0.85 accuracy has regressed against a baseline of 0.95, with avg latency 200, efficiency 0.01, and weights {'accuracy_weight': 0.5, 'latency_weight': 0.3, 'efficiency_weight': 0.2}.

{"compositeScore": 0.505, "regressionFlag": true}

You can use the `calculate_core_metrics` tool by providing an array of task results. It will return the F1 score along with accuracy, precision, and recall.

Related Connectors