Which AI models are available?

The NVIDIA API Catalog offers Llama 3.1 (8B, 70B, 405B), Mistral, CodeLlama, Gemma, Nemotron, and many more. Use the `list_models` tool to see all available models.

How do I get an NVIDIA API Key?

Sign up at [**build.nvidia.com**](https://build.nvidia.com), go to your account settings, and generate an API key. The Developer Program includes free inference credits.

Can I generate code in specific languages?

Yes! The `generate_code` tool lets you specify the programming language (Python, JavaScript, TypeScript, Java, etc.) for better results.

Are there usage limits on the free tier?

Yes, the NVIDIA Developer Program provides free inference credits. Once exhausted, you can upgrade to a paid plan for higher throughput. Check your usage dashboard at build.nvidia.com.

NVIDIA AI MCP Connector for Claude

A+

Access LLMs, embeddings, code generation, and reasoning via NVIDIA API Catalog.

9 tools Official Updated Jun 28, 2026 Official Vinkius Partner

More Details Connect to Claude

Connect NVIDIA AI to any AI agent and harness the power of GPU-accelerated foundation models — chat with Llama, generate embeddings, write code with CodeLlama, translate text, and perform complex reasoning through the NVIDIA API Catalog.

What you can do

Chat with LLMs — Access Llama 3.1, Mistral, Nemotron, and more via chat completions
Generate Embeddings — Create vector embeddings for search and clustering
Code Generation — Write code from natural language prompts using CodeLlama
Summarization — Condense long documents into concise summaries
Translation — Neural translation between dozens of languages
Text-to-SQL — Convert natural language questions into SQL queries
Sentiment Analysis — Analyze the emotional tone of text
Complex Reasoning — Ask questions to the 405B-parameter reasoning model

How it works

Subscribe to this server
Enter your NVIDIA API Key (from build.nvidia.com)
Start running AI models from Claude, Cursor, or any MCP-compatible client

Who is this for?

Developers — Prototype AI features without managing GPU infrastructure
Data Scientists — Generate embeddings and run NLP tasks at scale
Business Analysts — Use text-to-SQL to query databases with natural language

llmgpu-accelerationembeddingsmodel-inferencenatural-language-processingcode-generation