Markdown Semantic Chunker

Markdown Semantic Chunker MCP Connector for Claude

A+

A deterministic engine for splitting Markdown text into semantically coherent chunks based on header hierarchy and paragraph boundaries.

3 tools Official Updated Oct 1, 2026 Official Vinkius Partner

The Markdown Semantic Chunker MCP server provides a specialized engine for optimizing Retrieval-Augmented Generation (RAG) pipelines by preserving document structure during the chunking process. Unlike simple character-based splitters that often break sentences or headers mid-way, this tool uses a deterministic approach to identify header levels (# through ######) and group content accordingly. By maintaining a hierarchical path for every chunk, such as Intro > Setup > Step 1, it ensures that downstream LLMs receive the full structural context of each text segment. When a section exceeds the user-defined maxChunkSizeTokens, the engine intelligently splits at paragraph boundaries (double newlines) to maintain semantic integrity. The server includes specialized tools like generate_semantic_chunks for primary execution, get_header_structure for hierarchy previews, and calculate_markdown_token_density for analyzing text density.

markdownchunkingragsemantic-searchdocument-parsingnlp

3 tools expose this connector's capabilities to your AI agent.

generate_semantic_chunks

Generates semantic chunks from markdown content

get_header_structure

Extracts the header structure from markdown

calculate_markdown_token_density

Calculates the token density of markdown content

See how to talk to your AI agent using Markdown Semantic Chunker.

Split this markdown text into chunks with a maximum of 100 tokens: # Introduction This is the intro. ## Setup Step 1: Install dependencies.

The text was split into two chunks. Chunk 1: 'Introduction - This is the intro.' (Path: Introduction). Chunk 2: 'Setup - Step 1: Install dependencies.' (Path: Introduction > Setup).

Analyze the token density of this markdown block.

The analysis shows 2 distinct paragraphs and an estimated total of 45 tokens for the provided text block.

Show me all the headers in this document: # Title ## Section A ### Sub-section 1

The detected hierarchy is: ['Title', 'Title > Section A', 'Title > Section A > Sub-section 1'].

The tool assigns a hierarchical path to every chunk, representing its position in the document tree (e.g., Parent > Child). This ensures that even when text is split, the structural lineage remains attached to the content.

Related Connectors