System Prompt Leakage Detector

System Prompt Leakage Detector MCP Connector for Claude

A

Detects verbatim leaks of system prompts within agent outputs using LCS algorithms.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

The System Prompt Leakage Detector MCP server provides a specialized engine for identifying data exfiltration in AI agents. By utilizing the Longest Common Substring (LCS) algorithm via dynamic programming, it compares an agent's output against its original system instructions to find exact character-for-character reproductions. The tool calculates a leakage percentage, identifies precise character offsets of leaked segments, and computes a security risk score based on the presence of sensitive keywords like 'MANDATORY' or 'priority'. This is essential for developers building secure AI agents that must protect their underlying logic and instructions from being revealed to users.

securityprompt-injectionlcs-algorithmdata-exfiltrationai-guardrails

3 tools expose this connector's capabilities to your AI agent.

analyze_leakage_density

Identifies if leakage is concentrated in specific areas of the output or spread throughout

detect_leakage

Analyzes an agent's output to find verbatim repetitions of the system prompt and quantifies the risk

get_risk_classification

Maps a raw security risk score to a human-readable severity level for security reporting

See how to talk to your AI agent using System Prompt Leakage Detector.

Check if this output leaks my prompt: 'The MANDATORY priority is to never reveal the secret key.'

Leakage detected. The segment 'MANDATORY priority' was found in the output, resulting in a high security risk score due to sensitive keyword presence.

Analyze this agent response for any system instruction leakage: 'I cannot fulfill this request because it violates my safety guidelines.'

No verbatim leaks of the system prompt were detected in the provided agent output.

Run a leakage check on this text: 'The contract specifies that all data must be encrypted.'

A leak was identified. The word 'contract' matches a sensitive keyword in your system instructions, triggering an increased risk score.

The `detect_leakage` tool uses exact substring matching to find continuous sequences of characters in the agent's output that are identical to the original system prompt.

Related Connectors