Agent A/B Test Calculator

Agent A/B Test Calculator MCP Connector for Claude

A+

A deterministic statistical engine for evaluating performance differences between agent variants.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides a precise statistical engine to evaluate performance differences between different AI agent variants. It allows you to determine if a change in an agent's behavior resulted in a statistically significant improvement in conversion rates. Using tools like analyze_variant_performance, you can calculate p-values, confidence intervals, and relative lift while accounting for Bonferroni corrections in multi-variant tests. You can also use estimate_test_requirements to plan future experiments by calculating necessary sample sizes and test durations, or calculate_bayesian_probability to estimate the likelihood that one variant outperforms another using Beta distributions.

ab-testingstatisticsconversion-ratedata-analysismachine-learning

3 tools expose this connector's capabilities to your AI agent.

analyze_variant_performance

95 or 0.99), and the minimum detectable effect (MDE). Analyzes the performance of multiple agent variants to determine statistical significance

calculate_bayesian_probability

Calculates the Bayesian probability that variant B is better than variant A

estimate_test_requirements

Estimates the required sample size and duration for an A/B test

See how to talk to your AI agent using Agent A/B Test Calculator.

Analyze the performance of two variants: Variant A had 50 successes out of 1000 attempts, and Variant B had 70 successes out of 1000 attempts. Use a 95% confidence level.

Variant B shows a conversion rate of 7.0% compared to Variant A's 5.0%, resulting in a relative lift of 40.0%. This difference is statistically significant with a p-value of 0.012.

I expect a baseline conversion rate of 5% and want to detect a 2% MDE. My agent gets 100 attempts per day. How long will the test take with 95% confidence?

To detect a 2% MDE with 95% confidence, you will need a total sample size of approximately 3,850 attempts. At 100 attempts per day, the test will take about 39 days.

What is the probability that Variant B (10 successes, 100 total) is better than Variant A (8 successes, 100 total)?

The Bayesian probability that Variant B is better than Variant A is approximately 72.4%.

You can use the `analyze_variant_performance` tool. It calculates the p-value to determine if the observed difference in conversion rates is statistically significant or likely due to chance.

Related Connectors