Guardrail Violation Detector

Guardrail Violation Detector MCP Connector for Claude

A+

Analyzes AI agent outputs for safety and compliance violations.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides deterministic scoring for agent outputs to ensure safety and structural integrity. It evaluates responses against specific constraints like PII detection, toxicity, off-topic drift, length requirements, and format compliance. Using analyze_output_violations, agents can receive a composite violation score and severity level (Info, Warning, or Block) to decide whether to proceed with a response. It also provides access to the get_guardrail_configuration to retrieve active rules and get_historical_false_positive_rate to monitor system accuracy.

guardrailssafetycompliancepiitoxicity

3 tools expose this connector's capabilities to your AI agent.

analyze_output_violations

Performs a full audit of an agent response against a set of rules

get_guardrail_configuration

Retrieves the current active rules and their definitions

get_historical_false_positive_rate

Retrieves the accuracy metric for a specific guardrail rule

See how to talk to your AI agent using Guardrail Violation Detector.

Check if this response is safe: 'My phone number is 555-0199 and I am very angry!'

The response is flagged as a Block due to PII detection and high toxicity scores.

Analyze the compliance of a response that is 500 words long when the target was 100.

The response has a high length violation score due to the significant deviation from the target length.

What are the current active guardrail rules?

The active rules include PII detection, toxicity monitoring, off-topic detection, length constraints, and format validation.

Severity is based on the composite score: Info for scores below 0.2, Warning for 0.2 to 0.5, and Block for scores above 0.5. Additionally, any single rule exceeding 0.8 triggers a Block.

Related Connectors