System Prompt Leakage Detector

System Prompt Leakage Detector MCP Connector for Claude

A

Detects verbatim system prompt exfiltration using LCS algorithms.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

This MCP server provides specialized security tools to identify if an AI agent has leaked its foundational instructions. By using exact Longest Common Substring (LCS) matching, it compares agent outputs against the original system prompt to calculate leakage percentages and identify specific character offsets of leaked text. It includes tools like detect_leakage to quantify risk, analyze_leakage_density to find clusters of leaked text, and get_risk_classification to map scores to severity levels.

securityleakagelcsauditai-safety

3 tools expose this connector's capabilities to your AI agent.

analyze_leakage_density

Identifies if leakage is concentrated in specific areas of the output or spread throughout

detect_leakage

Analyzes an agent's output to find verbatim repetitions of the system prompt and quantifies the risk

get_risk_classification

Maps a raw security risk score to a human-readable severity level for security reporting

See how to talk to your AI agent using System Prompt Leakage Detector.

Check this output for system prompt leakage: 'The system instructions are: You are a helpful assistant.'

Leakage detected: 32% leakage with a High risk score due to sensitive instruction reproduction.

Analyze the density of these leaked segments: [0, 15], [40, 55] in a 100 character output.

The leakage is distributed with an average gap of 24 characters.

What is the severity for a risk score of 0.9?

The severity is Critical and immediate action is required.

The `detect_leakage` tool uses exact substring matching to find continuous sequences of characters in the agent's output that are identical to the original system prompt.

Related Connectors