Ollama

Ollama MCP Connector for Claude

A

Run LLM models via Ollama cloud API — generate completions, chat with multimodal models, create embeddings, and inspect model details from any AI agent.

12 tools Official Updated Oct 1, 2026 Official Vinkius Partner

Connect your Ollama API key to any AI agent and run large language models through natural conversation.

What you can do

  • Text Generation — Generate completions from any model (Gemma, GPT-OSS, Qwen, Llama) with support for images, structured outputs (JSON schema), system prompts, and thinking mode
  • Chat Conversations — Have multi-turn conversations with models, including multimodal inputs (vision), tool calling (function calling), and thinking output
  • Embeddings — Generate vector embeddings for semantic search, retrieval-augmented generation (RAG), and similarity matching
  • Model Discovery — List all available models, inspect detailed architecture info (parameters, capabilities, quantization, tokenizer settings)
  • Running Model Status — See which models are currently loaded in memory, their VRAM usage, and when they'll be unloaded
  • OpenAI Compatibility — Use OpenAI-compatible endpoints (/v1/chat/completions, /v1/completions, /v1/embeddings, /v1/responses) for drop-in compatibility with existing OpenAI SDK consumers
  • Version Diagnostics — Check the Ollama server version for compatibility and troubleshooting

How it works

  1. Subscribe to this server
  2. Enter your Ollama API key (create one at https://ollama.com/settings/keys)
  3. Start running models from Claude, Cursor, or any MCP-compatible client

Your AI acts as a gateway to Ollama's model library — generate text, have conversations, create embeddings, and explore models without writing a single line of code.

Who is this for?

  • Developers — prototype with different LLM models via natural language instead of reading API docs
  • AI Engineers — test model capabilities, generate embeddings, and compare outputs across models
  • Data Teams — create embeddings for RAG pipelines and semantic search applications
  • Product Teams — explore which models support vision, tools, and thinking capabilities before integrating
llmollamagemmagpt-ossqwenembeddingsopenai-compatiblemultimodal

12 tools expose this connector's capabilities to your AI agent.

generate_embeddings

Supports single text or array of texts, optional truncation for long inputs, and configurable output dimensions. Use for semantic search, retrieval, and RAG applications. Generate vector embeddings from text using a model

generate

Supports images (base64), structured outputs (JSON schema), system prompts, thinking mode, and model options (temperature, top_p, seed, etc.). Uses stream=false for a single complete response. Generate a text completion from a model

openai_chat_completions

Supports messages, tools, temperature, max_tokens, seed, response_format, vision (image_url), and reasoning_effort. Uses stream=false for a single complete response. Generate chat completions via the OpenAI-compatible endpoint

openai_embeddings

Compatible with the OpenAI embeddings API. Supports input as string or array of strings, optional encoding_format and dimensions. Generate embeddings via the OpenAI-compatible endpoint

openai_list_models

Returns model IDs, creation timestamps, and ownership info. Useful when using OpenAI SDKs that require this endpoint. List models via the OpenAI-compatible endpoint

openai_responses

Compatible with the OpenAI Responses API. Uses stream=false for a single complete response. Generate responses via the OpenAI Responses API endpoint

show_model_details

Get detailed information about a specific model

chat

Supports multimodal inputs (images as base64), tool calling, structured outputs (JSON schema), thinking mode, and model options. Pass the full messages array as JSON. Uses stream=false for a single complete response. Generate a chat message in a conversation

list_models

List all available models

list_running_models

Includes model name, size, VRAM usage, digest, context length, and expiry time (when the model will be unloaded). List models currently loaded in memory

openai_completions

Supports prompt, temperature, max_tokens, seed, stop sequences, and suffix (for fill-in-the-middle). Uses stream=false for a single complete response. Generate text completions via the OpenAI-compatible endpoint

get_version

Useful for diagnostics and compatibility checks. Get the Ollama server version

See how to talk to your AI agent using Ollama.

Generate a response from gemma4 for the prompt 'Why is the sky blue?'

Generating with gemma4... The sky appears blue because of a phenomenon called Rayleigh scattering. As sunlight enters Earth's atmosphere, the shorter blue wavelengths of light are scattered more than other colors by the molecules in the air. This scattered blue light comes to our eyes from all directions, making the sky appear blue.

List all available models and show me which ones support vision.

Found 12 models. Vision-capable models: gemma4 (8B, Q4_K_M, supports completion + vision), qwen3-vl:8b (8B, supports completion + vision + tools), and gpt-oss:20b (20B, supports completion + tools + thinking). The remaining 9 models support text completion only. Would you like detailed architecture info for any specific model?

Create embeddings for the text 'The quick brown fox jumps over the lazy dog' using nomic-embed-text.

Generated 768-dimensional embeddings for the input text using nomic-embed-text. The embedding vector has been created successfully. Processing took 8 tokens in the prompt evaluation. Would you like me to generate embeddings for multiple texts at once, or adjust the output dimensions?

Visit the [Ollama API keys page](https://ollama.com/settings/keys) to create or manage your API key. Once created, copy it and paste it into the API key field in the setup wizard. The key is used as a Bearer token in the Authorization header for all requests to https://ollama.com/api. API keys don't currently expire, but you can revoke them at any time.

Related Connectors