AI Agent Tools With Open-Weight Model Support
24 of the 94 tools in our index document real support for at least one open-weight or locally-run model (Llama, Mistral, DeepSeek, Qwen and others, typically via Ollama, LM Studio or a LiteLLM-style router) alongside the closed frontier APIs: a real answer to the single-vendor lock-in question, spanning frameworks, coding agents, agent platforms, voice agents and research agents.
On this page
Being "model-agnostic" and supporting an open-weight model are different claims, and this collection only counts the second one: a tool is listed here because its own published listing names a specific open-weight family — Llama, Mistral, DeepSeek, Qwen, GLM or Kimi among them — or explicit local execution through a runtime like Ollama or LM Studio. 24 of the 94 tools in this index clear that bar.
Several tools that look like candidates fail on exactly that distinction rather than on capability. [LangChain](/langchain) describes its model support as "Model-agnostic (any LLM)", [Vapi](/vapi) as "Model-agnostic (bring your own)", [Retell AI](/retell-ai) as "Model-agnostic (choose your LLM)", and [LangGraph](/langgraph) names "OpenAI, Anthropic, Google, and more" — three closed providers and an "and more". Each of the four may well run an open-weight model; nothing any of them publishes says so, so none of them is counted here.
By category: 5 of 14 [agent frameworks](/category/frameworks) (Microsoft Agent Framework, Pydantic AI, Agno, Google ADK, CrewAI), 8 of 25 [coding agents](/category/coding-agents) (Cline, Aider, Juggler, OpenAI Codex CLI, OpenCode, Zot, Zed and Junie, which JetBrains' own docs confirm can point at a local Ollama or LiteLLM runtime alongside its frontier-model options), 6 of 16 [agent platforms](/category/agent-platforms) (n8n, Dify, Botpress, OpenClaw, Vecbase), 3 of 12 [voice agents](/category/voice-agents) (ElevenLabs Conversational AI, Deepgram Voice Agent API, Cartesia) and 3 of 11 [research agents](/category/research-agents) (GPT Researcher, Perplexity and Genspark). Genspark is the one entry where the open-weight model is the vendor's choice rather than yours: its Super Agent blends nine LLMs, DeepSeek among them, and you do not pick which one answers. Everywhere else on this page the open weights are something you select, host or point the tool at.
Our remaining categories — [support agents](/category/support-agents), [sales & marketing agents](/category/sales-marketing-agents) and the newer [agent tools](/category/agent-tools) — have no entry here at all. Every listing in all three is closed-model and cloud-only by its own published record.
Why this collection
A tool earns a place here only if its own Published, sourced listing names at least one specific open-weight model family (Llama, Mistral, DeepSeek, Qwen, GLM, Kimi or similar) or explicit local-model execution (Ollama, LM Studio, "local model") somewhere in its `whatIs`, `whatItDoes`, `keyFeatures`, `integrations` or `attributes.modelLLM` fields: a generic "model-agnostic" or "bring your own model" claim with no open-weight option ever named does not qualify on its own. Verified directly against each listing's own text, not inferred from the `open-models` tag alone in either direction: a tagged listing that names nothing specific is excluded, and an untagged listing whose own text clears the bar is included.
24 agents in this collection
Best for .NET or Python teams already invested in Azure or Microsoft Foundry who want one supported agent SDK, an optional managed hosting path, and a documented migration off AutoGen or Semantic Kernel.
Builds an agent from Azure OpenAI, OpenAI, Microsoft Foundry, Anthropic Claude, local models via Ollama or a custom provider — genuinely model-agnostic despite the Microsoft branding.
Best for Python teams building type-safe agents who also want realtime voice, image generation and embeddings in one SDK, with optional durable execution via Temporal, DBOS, Prefect or Restate.
Model-agnostic across OpenAI, Anthropic, Gemini, Mistral, Cohere, DeepSeek, Grok, Bedrock and local models via Ollama — one of the widest named provider lists in this collection.
Best for Python developers who want an Apache 2.0 SDK with memory, knowledge, learning, guardrails and a Control Plane they run inside their own cloud account, and who are comfortable running Docker or a cloud deploy template themselves.
Interchanges 20+ LLM providers including Ollama for local models without rewriting agent code, and a free AgentOS Control Plane runs entirely on your own machine for local development.
Best for Teams that want Google's own official, actively-developed agent SDK, especially polyglot teams (Python/Go/Java/TypeScript/Kotlin) or anyone building multi-agent systems that need to interoperate across vendors via the A2A protocol.
Optimized for Gemini but genuinely model-agnostic through a LiteLLM integration — OpenAI, Anthropic Claude, Cohere, local Ollama models and 100+ other providers are all reachable the same way.
Best for Developers who want to stand up a role-based, collaborating multi-agent system quickly in Python, with a stable MIT-licensed core and a hosted build and runtime for production.
Its integrations list routes through LiteLLM to providers beyond the OpenAI, Gemini and Hugging Face options it names directly, reaching local and self-hosted models via Ollama.
Best for Ops and automation teams who want to build agentic workflows visually while keeping data on self-hosted infrastructure.
Its native AI-agent and LLM nodes list OpenAI, Anthropic, Ollama and more — wire a local Ollama model into the same visual canvas as any hosted provider.
Best for Developers and teams who want an open-source, self-hostable platform to build RAG apps and agents with full data control.
Integrates hundreds of LLMs from dozens of providers, explicitly including open and local models alongside OpenAI, Anthropic and Google, so a node can be swapped to a self-hosted model without rebuilding the flow.
Best for Engineering-led teams building a multi-channel AI support agent who want an open-source, self-hostable core with an optional managed cloud layer.
Model-agnostic across OpenAI, Anthropic Claude, Google Gemini and open-weight models, per its own published spec.
Best for Developers and teams who want a transparent, open-source, model-agnostic agent they fully control and pay for by usage, with optional flat-rate access to open-weight models.
Bring-your-own-key across Anthropic, OpenAI, DeepSeek, xAI, Mistral, Cerebras or a local model via Ollama, plus a paid ClinePass tier for discounted access to a curated set of open-weight models (Qwen, DeepSeek, Kimi, GLM and others).
Best for Developers who want free, model-agnostic AI pair programming in their terminal
Bring-your-own-key and model-agnostic across Anthropic, OpenAI, DeepSeek, local models and effectively any LLM via Ollama — free either way, since you pay the model provider directly.
Best for Developers who want hands-on visibility into what an AI coding agent is doing to their codebase and accept early-stage, self-hosted, open-source software.
Drives Claude, OpenAI, Gemini, Ollama, OpenRouter or DeepSeek models through a swappable JavaScript plugin system — free, open-source and model-agnostic with no usage cap of its own.
Best for Developers who already pay for a ChatGPT plan and want a terminal-native agent with granular, inspectable sandbox/permission controls across CLI, desktop and IDE surfaces.
Ships built-in support for local/self-hosted models via Ollama and LM Studio alongside Amazon Bedrock and OpenAI's own models — any other provider needs a manual config.toml entry.
Best for Developers who want a genuinely free, fully open-source coding agent that works with any model provider, including their own local Ollama models, without being locked into one vendor's terminal tool.
Bring any of 75+ model providers, including your own local Ollama models, reachable through Models.dev or wired in directly via the /connect command.
Best for Developers who want a fast, free, Go-native coding-agent harness with real session-branching control and are comfortable running fast-moving, self-described "beta forever" software from an independent maintainer.
Connects to 30+ model providers — Anthropic, OpenAI/Codex, Gemini, Kimi, DeepSeek, GitHub Copilot, Bedrock, OpenRouter and local Ollama models — through one static Go binary.
Best for Developers who want a fast, open-source editor and the freedom to run Claude Code, Codex, Gemini CLI or other vendors' own agents natively via the Agent Client Protocol Zed created, instead of being locked into one editor's built-in agent.
Connects to ten-plus model providers including Mistral, DeepSeek, Ollama and LM Studio, and ships its own first-party open-weight Zeta2 model for inline edit prediction, included in every plan.
Best for Developers and product teams who want ElevenLabs' voice quality and cloning as the foundation of their agent, and are comfortable assembling the rest of the stack themselves.
An OpenAI-compatible custom endpoint opens the door to open-source models like Llama 3.3 or DeepSeek R1 via hosts such as Cloudflare Workers AI, Groq or Together AI, alongside the built-in OpenAI, Anthropic and Google options.
Best for Engineering teams that want a managed speech pipeline with real LLM-provider choice, and regulated enterprises that need a self-hosted or VPC deployment option a typical closed-source voice-agent vendor does not offer.
Its LLM layer spans OpenAI, Anthropic, Google, NVIDIA, Groq and Amazon Bedrock — more than 30 specific models, including open-weight options such as Groq's gpt-oss-20b and NVIDIA's Nemotron-3-Nano alongside closed frontier models.
Best for Engineering teams that want an open-source voice-agent SDK with real LLM choice, built on speech models several rival platforms in this category already license as a component.
The Line SDK connects to more than 100 LLM providers via LiteLLM, including open-weight models by name, rather than locking a buyer into Cartesia's own reasoning layer.
Best for Developers and AI engineers who want a free, self-hosted research agent they can embed via Python, REST or MCP and run against their own LLM and search keys.
Model-agnostic — drive it with OpenAI, Anthropic, Google, Groq or local Ollama models, pairing whichever LLM you choose with a search back-end like Tavily or Bing.
Best for Developers and teams already standardized on a JetBrains IDE who want an agent that shares the same debugger, project index, and version control already built into their editor.
JetBrains' own docs confirm Junie can point at a local runtime via Ollama or LiteLLM alongside its OpenAI/Anthropic/Google frontier-model options — a genuine open-weight path, not just "model-agnostic" marketing language.
Best for Developers and self-hosters who already run their own infrastructure and want one personal AI agent that reaches WhatsApp, Telegram, Slack, Discord, Signal and iMessage instead of a separate app.
Routes requests through 64 model and media providers including local runtimes — Ollama and LM Studio — alongside Anthropic, Google and OpenAI, plus gateways such as Amazon Bedrock and OpenRouter.
Best for Small teams or solo operators who want pre-built AI agent roles, shared files, and OAuth-driven workflows under one credit-based subscription.
Agents switch between multiple underlying models — Claude, GPT and Gemini beside the open-weight Kimi and DeepSeek — with transparent per-token billing on whichever one you pick.
Best for Researchers, analysts and knowledge workers who want fast, sourced answers, plus an autonomous research mode and an agentic browser.
Names z.ai's GLM and Moonshot AI's Kimi, both added in 2026, alongside its own Sonar models and GPT, Claude and Gemini — the open-weight options are selectable, not just routed to behind the scenes.
Best for A consultant, marketer, founder or operations lead who wants research, files, code, slides and media handled in one web workspace.
Its Super Agent blends nine specialized LLMs including DeepSeek — the one entry here where the open-weight model is chosen by the vendor's router rather than by you.