Research

Research reportai-agents

Which AI Agents Won't Tell You What LLM Powers Them?

We checked all 44 tools in our index for what LLM sits under the hood. 10 disclose nothing at all — and the pattern splits almost perfectly by category.

“Powered by GPT-5” and “Claude inside” are becoming marketing lines in this space, but plenty of vendors say nothing at all about which model is doing the work. That’s not a small detail — it decides whether you can control cost, satisfy a compliance team, or switch providers if a model gets worse or a vendor’s pricing changes. We went through our own 44-listing index and classified every tool by how much it actually tells you about its underlying LLM. The headline: 10 of 44 (23%) disclose nothing beyond “proprietary” or “managed, multi-model” — and which tools stay silent tracks almost exactly with whether the product is sold as software you configure or a fully-managed team you hire.

How we classified each tool

Every listing in our index carries a modelLLM field, researched and sourced the same way as pricing or API availability. We sorted all 44 current values into five tiers, from most to least transparent:

  1. Agnostic / BYOK — you can bring or choose any LLM, including through an open routing layer (LiteLLM, OpenRouter, a built-in Ollama/local-model option).
  2. Closed menu (multi) — the vendor names 2 or more specific frontier models you can pick between, but there’s no open routing or bring-your-own-key option.
  3. Single-vendor locked — the product runs on exactly one named vendor’s models, full stop.
  4. Proprietary, named — the vendor built and named its own in-house model(s). Not third-party, but not a mystery either.
  5. Not disclosed — “proprietary,” “multi-model,” or “managed” with no further detail anywhere in the vendor’s own materials.

The full breakdown

Tool Category Transparency What the vendor says
LangChain Agent frameworks Agnostic / BYOK Model-agnostic (any LLM)
CrewAI Agent frameworks Agnostic / BYOK Model-agnostic (via LiteLLM)
Microsoft AutoGen Agent frameworks Agnostic / BYOK Model-agnostic (OpenAI/Azure first-class)
LlamaIndex Agent frameworks Agnostic / BYOK Model-agnostic
OpenAI Agents SDK Agent frameworks Agnostic / BYOK OpenAI-first (100+ via LiteLLM)
Pydantic AI Agent frameworks Agnostic / BYOK Model-agnostic
Cline Coding agents Agnostic / BYOK Model-agnostic (BYO key)
Aider Coding agents Agnostic / BYOK Model-agnostic (BYO key)
Juggler Coding agents Agnostic / BYOK Model-agnostic (Claude, OpenAI, Gemini, Ollama, OpenRouter, DeepSeek)
OpenAI Codex CLI Coding agents Agnostic / BYOK OpenAI GPT-5.6 family, plus built-in Ollama, LM Studio, Amazon Bedrock
OpenCode Coding agents Agnostic / BYOK 75+ providers via Models.dev, plus OpenCode Zen/Go hosted options
Zot Coding agents Agnostic / BYOK 30+ providers (Anthropic, OpenAI/Codex, Gemini, Kimi, DeepSeek, Bedrock, Azure, OpenRouter, Ollama…)
OpenHands Coding agents Agnostic / BYOK Bring your own API key, or use OpenHands’ own models at cost (commonly Claude or GPT)
Cursor Coding agents Closed menu Claude, GPT & Gemini (per-request)
GitHub Copilot Coding agents Closed menu GPT, Claude & Gemini (selectable)
Devin Coding agents Closed menu Claude, GPT & Gemini + Cognition’s own SWE models
Windsurf Coding agents Closed menu SWE models + Claude, GPT & Gemini
Replit Agent Coding agents Closed menu Claude (primary) & other frontier models
Claude Code Coding agents Single-vendor locked Claude (Opus & Sonnet) only
Bolt.new Coding agents Single-vendor locked Claude (Sonnet 4.6 default, Opus 4.7, Haiku 4.5) only
Vapi Voice agents Agnostic / BYOK Model-agnostic (bring your own)
Retell AI Voice agents Agnostic / BYOK Model-agnostic (choose your LLM)
ElevenLabs Conversational AI Voice agents Agnostic / BYOK ElevenLabs-hosted, or bring your own OpenAI-compatible endpoint
LiveKit Agents Voice agents Agnostic / BYOK Bring your own, or LiveKit-hosted inference
Bland AI Voice agents Not disclosed Runs its own self-hosted model stack — architecture disclosed, model identity is not
n8n Agent platforms Agnostic / BYOK Model-agnostic (OpenAI, Anthropic, Ollama…)
Dify Agent platforms Agnostic / BYOK Model-agnostic (100s of LLMs)
Lindy Agent platforms Closed menu Multi-model
Relevance AI Agent platforms Closed menu Claude, GPT & Gemini (router)
GPT Researcher Research agents Agnostic / BYOK Model-agnostic (OpenAI, Anthropic, Google, Ollama…)
Perplexity Research agents Closed menu Sonar (own), GPT, Claude, Gemini, GLM & Kimi
Manus Research agents Closed menu Multi-model (frontier LLMs)
Elicit Research agents Not disclosed Multi-model backend, not publicly disclosed
GC AI Research agents Not disclosed Multi-model backend, not publicly disclosed
Clay Sales & marketing agents Closed menu Claygent — GPT & Claude (selectable)
11x Sales & marketing agents Not disclosed Proprietary
Artisan Sales & marketing agents Not disclosed Proprietary
Qualified Sales & marketing agents Not disclosed Proprietary (managed), not publicly disclosed
Rox Sales & marketing agents Not disclosed Not publicly disclosed (managed, multi-model stack)
Sierra Support agents Closed menu Multiple frontier models (managed)
Intercom Fin Support agents Proprietary, named Runs on Intercom’s own named Apex 1.0 / Apex Flash models
Decagon Support agents Not disclosed Proprietary (managed, multi-model)
Crescendo Support agents Not disclosed Not publicly disclosed (managed, multi-model stack)
Cresta Support agents Not disclosed Not publicly disclosed (managed, multi-model stack)

The category split is the real finding

Looking at the 44 tools as one list, transparency looks like a coin flip: 20 agnostic, 11 on a closed menu, 2 locked, 1 named-proprietary, 10 silent. Split by category, it isn’t a coin flip at all:

Category Agnostic / BYOK Not disclosed Sample
Agent frameworks 6 of 6 (100%) 0 of 6 (0%) n=6
Voice agents 4 of 5 (80%) 1 of 5 (20%) n=5
Coding agents 7 of 14 (50%) 0 of 14 (0%) n=14
Agent platforms 2 of 4 (50%) 0 of 4 (0%) n=4
Research agents 1 of 5 (20%) 2 of 5 (40%) n=5
Sales & marketing agents 0 of 5 (0%) 4 of 5 (80%) n=5
Support agents 0 of 5 (0%) 3 of 5 (60%) n=5

(Sierra is dual-listed under support agents and agent platforms; it’s counted once, under support agents, above.)

Agent frameworks — libraries you install and wire up yourself — are unanimously agnostic; that’s close to the point of a framework. Coding agents split roughly in half, but critically, all 14 tell you something about their models even when they won’t let you swap them (Claude Code and Bolt.new are explicit about being Claude-only, not silent about it). Support agents and sales & marketing agents are the mirror image: not one of the 10 tools across both categories is model-agnostic, and 7 of them disclose nothing beyond “proprietary” or “managed.” Both categories are sold the same way — as an outsourced, fully-managed team rather than software you configure — and that sales model appears to come with an unwritten rule: the buyer doesn’t get to know, or ask, what’s running underneath.

Why this is a real buying question, not trivia

Four concrete reasons the model behind an agent matters even if you never touch a system prompt:

  • Cost control. An agnostic or BYOK tool lets you route to a cheaper model for high-volume, low-stakes tasks and a frontier model only where it earns its keep. A vendor charging a flat per-seat or per-resolution price on an undisclosed model gives you no way to reason about their margin or your own unit economics.
  • Compliance and data residency. If you can’t find out which model — or whose infrastructure — processes your data, you can’t answer a security questionnaire about it. Bland AI is the one exception worth naming here: it won’t say which model it runs, but it’s explicit that the model runs on its own infrastructure rather than routing through a third-party provider, which is itself a meaningful, disclosed compliance fact even without a model name.
  • Quality and behavior drift. A vendor can swap the model behind a “managed” product at any time with no changelog entry a buyer would see — you’d only notice output quality or behavior change. Agentic tools with a documented model (or a choice of one) at least tell you what changed if a vendor announces an upgrade.
  • Switching leverage. If a vendor doubles its price or a model’s outputs regress, an agnostic tool lets you point at a different model tomorrow. A locked or undisclosed one leaves you with no lever but to renegotiate or migrate the entire product.

None of this means “not disclosed” is automatically a red flag — an enterprise buyer evaluating Crescendo or Decagon is typically buying an outcome (resolved tickets, booked meetings) and delegating the infrastructure decision on purpose. It does mean that buyer is trusting the vendor’s model choices sight unseen, which is worth naming explicitly during procurement rather than discovering later.

Proprietary doesn’t have to mean silent

Intercom Fin is the one listing in our “proprietary” group that isn’t actually opaque: Intercom names its in-house models directly — Apex 1.0 for general resolution, Apex Flash tuned for low-latency voice — in its own product materials. That’s a useful contrast with the other nine “not disclosed” tools, which give a category (support, sales, research) but never a model name, whether their own or a third party’s. Being closed-source and proprietary is a legitimate product choice; refusing to name what you built is a separate, stricter kind of opacity — and only one vendor in our proprietary group draws that distinction in its own favor.

What this census doesn’t tell you

“Not disclosed” here means not stated anywhere in the vendor’s own public materials as of this research pass — not that the tool’s output is worse, or that the vendor is hiding something disreputable. Several of these products (Decagon, Crescendo, Cresta) are well-funded, enterprise-vetted platforms with strong compliance credentials elsewhere (SOC 2, HIPAA, GDPR); model disclosure is simply a fact they’ve chosen not to publish, possibly for competitive reasons. This is also a snapshot: a vendor could publish its model stack the day after this runs, and several products (Perplexity’s Sonar, Windsurf’s SWE models) already show that a company can name its own model and offer frontier alternatives — the two aren’t mutually exclusive.

Methodology

Every row above comes directly from our own listings’ sourced, gated research — not a new survey, and nothing inferred from marketing copy alone. Category shares are computed against the current Published corpus (44 listings, 7 categories, one dual-tagged) and will shift as we add or re-verify listings; this article isn’t on the update cadence our State of AI Agents 2026 report is, so treat the date above as when it was last checked against the live index. If you’re evaluating a specific tool, verify directly with the vendor before you rely on a classification here — model support is one of the faster-moving facts in this space.

Get the next report

New agents rankings and fresh data reports. One short email, straight to your inbox. One-click unsubscribe.