Ollama MCP Connector for Claude
ARun LLM models via Ollama cloud API — generate completions, chat with multimodal models, create embeddings, and inspect model details from any AI agent.
Connect your Ollama API key to any AI agent and run large language models through natural conversation.
What you can do
- Text Generation — Generate completions from any model (Gemma, GPT-OSS, Qwen, Llama) with support for images, structured outputs (JSON schema), system prompts, and thinking mode
- Chat Conversations — Have multi-turn conversations with models, including multimodal inputs (vision), tool calling (function calling), and thinking output
- Embeddings — Generate vector embeddings for semantic search, retrieval-augmented generation (RAG), and similarity matching
- Model Discovery — List all available models, inspect detailed architecture info (parameters, capabilities, quantization, tokenizer settings)
- Running Model Status — See which models are currently loaded in memory, their VRAM usage, and when they'll be unloaded
- OpenAI Compatibility — Use OpenAI-compatible endpoints (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/responses) for drop-in compatibility with existing OpenAI SDK consumers
- Version Diagnostics — Check the Ollama server version for compatibility and troubleshooting
How it works
- Subscribe to this server
- Enter your Ollama API key (create one at https://ollama.com/settings/keys)
- Start running models from Claude, Cursor, or any MCP-compatible client
Your AI acts as a gateway to Ollama's model library — generate text, have conversations, create embeddings, and explore models without writing a single line of code.
Who is this for?
- Developers — prototype with different LLM models via natural language instead of reading API docs
- AI Engineers — test model capabilities, generate embeddings, and compare outputs across models
- Data Teams — create embeddings for RAG pipelines and semantic search applications
- Product Teams — explore which models support vision, tools, and thinking capabilities before integrating
Related Connectors
Baidu Qianfan MCP
Orchestrate Baidu Qianfan AI models — manage chat completions, embeddings, and prompt templates directly from any AI agent.
SambaNova (AI Inference) MCP
High-speed AI inference for Llama 3, DeepSeek, and MiniMax models via SambaNova's ultra-fast SN40L chips.
Voyage AI (AI Embeddings API) MCP
Generate high-quality text, multimodal, and contextualized embeddings, plus high-precision reranking for RAG workflows.
Hugging Face LLM MCP
Connect Hugging Face LLM to any AI agent via MCP.