How to Choose an AI Research Agent: 11 Compared
All 11 research agents in our index look like rivals, but each is built for a different job: web answers, literature, legal, raw web data, or your own data.
On this page
Scope · 13 topics
- Research agents
- Decision guide
- Perplexity
- Manus
- Genspark
- GPT Researcher
- Webhound
- Elicit
- Undermind
- GC AI
- Microsoft Fabric Data Agent
- Julius AI
- TinyFish
Perplexity, Manus, Genspark, GPT Researcher, Webhound, Elicit, Undermind, GC AI, Microsoft Fabric Data Agent, Julius AI and TinyFish all get called "research agents," and a flat ranked list makes them look like eleven ways to buy the same thing. Read their own full researched listings side by side and that's not true: each is built to answer a different question, for a different buyer. Perplexity answers everyday questions off the live web and layers on an autonomous research mode and an agentic browser. Manus isn't really a research tool first: it's a general-purpose autonomous agent that happens to be good at research alongside building decks, sites and small apps. Genspark takes a related but distinct bet on the same broad-agent idea: its Super Agent blends nine LLMs with cross-model verification and adds real outbound phone calls and a personal-assistant layer to Manus's research-and-build scope, but, unlike Manus, publishes no developer API. GPT Researcher is an open-source pipeline you self-host and wire to your own model and search keys, with no vendor and no subscription. Webhound researches the same open web as Perplexity and GPT Researcher, but priced and delivered differently from both: a dollar budget instead of a subscription or self-hosting, and a choice of a cited Report or a structured Dataset, triggerable directly from Claude, Codex or Cursor via a native MCP server. Elicit and Undermind both scope narrowly and deliberately to scientific and academic literature, but for opposite ends of a literature review: Elicit screens a candidate set of papers you already have against inclusion criteria; Undermind tries to find the papers you didn't know to look for in the first place. GC AI is scoped just as narrowly, but to a different body of knowledge entirely: US case law, statutes, regulations and agency guidance, for in-house legal teams rather than researchers or developers. Microsoft Fabric Data Agent researches nothing external at all: it turns a plain-English question into a governed SQL, DAX or KQL query over your own Fabric lakehouse, warehouse or Power BI model, and it is one of only two here scoped to private data rather than any public corpus. Julius AI is the other, and it is aimed at a completely different buyer: no data platform required, just a spreadsheet you upload or a Snowflake, BigQuery or Postgres connection, asked in plain language and answered as a chart, a report or a finished presentation. TinyFish doesn't produce a research answer at all: it's a web-data API bundle (Search, Fetch, Agent, Browser) that extracts structured data off live pages or completes a multi-step browser task, callable directly from Claude, Cursor or n8n, and priced per request through a prepaid Wallet rather than a subscription; the other ten answer a question, TinyFish fetches the raw material a question would be answered from. Elicit's own listing even names two of the others directly: it says skip it for "a general web-search answer engine (see Perplexity) or a broad autonomous agent that automates open-web tasks and deliverables beyond published literature (see Manus)": a clear signal that these eleven aren't really competing for the same buyer.
The eleven, at a glance
Every figure below is pulled from each tool's own researched listing, verified 23 July–18 August 2026.
| Agent | What it actually is | Entry price | Open source? | Built for |
|---|---|---|---|---|
| Perplexity | AI answer engine with an autonomous Deep Research mode and the agentic Comet browser | Free → $20/mo Pro → $200/mo Max → custom Enterprise | No (closed) | Fast, sourced answers to everyday questions, plus autonomous web research on demand |
| Manus | General-purpose autonomous agent that plans and executes multi-step tasks in its own cloud computer | Free (~300 daily credits) → Pro from $20/mo (4,000–40,000 credits) → Team from $20/seat/mo | No (closed) | Hands-off research plus finished deliverables: reports, decks, small deployed apps |
| Genspark | Best-funded general agent, a "Super Agent" blending 9 LLMs with cross-model verification across research, slides, sheets and real phone calls | Free (~100 daily credits) → Plus $24.99/mo (10,000 credits) → Pro $249.99/mo (125,000 credits) → custom Team/Enterprise | No (closed) | One broad research-to-deliverable workspace, including outbound phone calls, with no developer API |
| GPT Researcher | Open-source, self-hosted planner-executor research agent you run yourself | Free (Apache-2.0): you pay only your own LLM + search API usage | Yes (Apache-2.0) | Developers who want a research pipeline they fully control, with no subscription and no vendor |
| Webhound | Hosted, agent-triggered research agent priced per dollar of compute, delivering a cited Report or a structured Dataset | Pay-as-you-go: roughly $1 per 15 minutes (one-time $5 signup credit) | No (closed) | Developers who want research callable from Claude, Codex or Cursor via MCP, priced transparently rather than by subscription or credits |
| Elicit | Research agent purpose-built for scientific and academic literature (138M+ papers) | Free Basic → $49/mo Pro → $169/mo Scale → custom Enterprise | No (closed) | Literature reviews, formal systematic reviews and structured data extraction from published research |
| Undermind | AI literature-discovery agent that reads full papers and follows citation trails to find work a keyword search misses | Free → $16/mo Pro (billed annually) → $15/person/mo Team → custom Enterprise | No (closed) | Exploratory literature discovery and novelty checks on niche or cross-disciplinary questions |
| GC AI | Research agent for in-house legal teams, running parallel cross-checked research over a 13M+-opinion US case-law and statute corpus | $500/mo Individual (14-day trial) → Team on request → custom Enterprise | No (closed) | In-house counsel researching case law, statutes and regulations without engaging outside counsel for a first pass |
| Microsoft Fabric Data Agent | Natural-language-to-SQL/DAX/KQL agent built into Microsoft Fabric, governed by the requesting user's own permissions and Microsoft Purview policy | No standalone price: requires an existing paid Fabric F2+ or Power BI Premium P1+ capacity (independent estimate: ~$263/mo minimum) | No (closed) | Teams already on Fabric/Power BI who want governed natural-language Q&A over their own warehouse or BI data |
| Julius AI | Chat-based data-analysis agent over files, databases and warehouses you connect yourself, returning charts, reports and presentations | Free → $20/mo Plus ($16 billed annually) → $45/mo Pro → $200/mo Max → $450/mo Business | No (closed) | Knowledge workers and small teams who want to question their own spreadsheets or warehouse without writing SQL or buying a data platform |
| TinyFish | Web-data API bundle (Search, Fetch, Agent, Browser) under one key, extracting structured data or completing multi-step browser tasks rather than answering a question | Free (Search & Fetch, no card) → pay-as-you-go Wallet: $0.016/step Agent, $0.002/min Browser → custom Enterprise | Partial (Agent component only, not independently confirmed) | Developers pulling or automating structured data off the live web at scale, callable by API or from Claude, Cursor or n8n |
The decision: which job do you actually need done?
All eleven carry a full researched verdict on their own listing, including who each one is explicitly not for. Read together, the honest decision tree runs on the shape of the job, not feature checklists:
- You want fast, sourced answers to everyday questions, and an autonomous research mode when a question needs more than one search → Perplexity. Its Deep Research mode runs many searches and compiles a cited report on its own, and the Comet browser can carry out multi-step tasks on the live web for you. Watch out for: it's fundamentally a Q&A tool with agentic layers on top, not a general task agent, and like any LLM search product it can still attach a citation that doesn't fully support the claim. The unlimited Max tier is also expensive at $200/month.
- You want to hand off a whole multi-step task (not just research, but a finished report, slide deck or small app) and can tolerate uneven reliability on complex work → Manus. It genuinely plans and executes long task chains in its own cloud computer, browsing, coding and producing a finished deliverable with little supervision. Watch out for: independent testers found it degrades on long, branching workflows and has produced fabricated "verification" output that looked real, its credit-based billing can burn hundreds to over a thousand credits on one complex job, and its corporate backing has been unsettled since Meta's roughly $2B acquisition of its parent company (closed December 2025) was ordered fully unwound by China's regulator in April 2026.
- You want that same broad, hands-off scope but with cross-model verification, real outbound phone calls and a personal-assistant layer, from the best-funded vendor in this category → Genspark. Its Super Agent blends nine LLMs (GPT, Claude, Gemini, DeepSeek and others), cross-checking outputs between them, and adds Sparkpages research reports, AI Slides/Sheets, Call for Me phone calls and a Genspark Browser autopilot mode, backed by $535M raised and roughly $250M in annualized revenue. Watch out for: no source found, vendor or independent, documents a public developer API, unlike Manus, so there's no confirmed way to build it into your own software; independent reviews also flag a session cap that undermines "unlimited" chat and credits that expire unused rather than rolling over.
- You want full control (self-hosted, model-agnostic, zero subscription) and you're comfortable installing and running a developer tool → GPT Researcher. It's free and open-source (Apache 2.0), you bring your own LLM and search keys, and its Deep Research mode can recursively explore a topic across many branches for a few cents to about a dollar in API spend per run. Watch out for: there's no polished hosted product, you manage your own setup, and report quality depends entirely on the model and search backend you choose.
- You want a hosted research agent your own AI coding tool can call directly, priced per dollar of compute rather than a subscription or a credit meter → Webhound. Its native MCP server lets Claude, Codex or Cursor trigger a run and read the output as a tool call, and it returns either a cited Report or a structured Dataset with a source trail on every value. Watch out for: it's a two-person, YC-backed team that has already changed its own pricing and interface more than once since a mid-2025 free-tool launch, and its own blog states that building trust in AI-generated research remains an unsolved problem for them, not a solved one.
- You already have a candidate set of papers (or a formal review protocol) and need to screen them against inclusion criteria at scale, with a documented, auditable process → Elicit. It searches 138M+ papers, and its Systematic Review workflow automates screening against inclusion criteria at real scale (up to 40,000 papers on Enterprise), with every generated claim carrying a sentence-level citation back to the source paper. Watch out for: independent reviewers say it can miss papers a manual database search would catch (especially very recent ones), Elicit's own documentation self-discloses hallucination risk and tells users to verify findings against the source, and Pro pricing ($49/mo) is steep next to narrower single-purpose competitors if all you need is quick paper lookups.
- You're at the earlier, exploratory stage: you don't yet know what you don't know, and want to catch niche or cross-disciplinary papers a keyword search would miss → Undermind. It reads and evaluates hundreds of full-text papers and follows citation trails, then reports a "discovery curve" estimate of how complete that search was, independently reviewed (Journal of the Canadian Health Libraries Association, Aug 2025) as genuinely finding papers others miss. Watch out for: the same review calls its inability to produce a reproducible search strategy its "Achilles heel" (a real gap for a formal systematic review); searches take 8–10 minutes, and there's no API today.
- You're in-house counsel or legal ops, and the question spans case law, a statute or agency guidance → GC AI. Its Research Agent runs specialized agents in parallel across jurisdictions and courts over a 13M+-opinion corpus, then cross-checks and reconciles the findings into one cited analysis, the only listing in this category scoped to legal research at all. Watch out for: on the cheapest $500/month Individual plan, US Case Law research (the feature that makes this a "research agent" rather than a general legal chatbot) is a paid add-on rather than included by default, and the public API, though self-serve to access, is still in beta with conservative rate limits ahead of general availability.
- The "research" you need done is a question about your own company's data, not anything on the public internet → Microsoft Fabric Data Agent. It turns a plain-English question into a governed SQL, DAX or KQL query over your own Fabric lakehouse, warehouse or Power BI model, with every query scoped automatically to the requesting user's own permissions and Microsoft Purview policy: the governed choice of the two here that research nothing external at all, and the only one that inherits your existing permission model instead of asking you to connect a database to a chat product. Watch out for: you cannot change the underlying model, it only reads structured data (no PDFs, no non-English questions), and there is no way to try it without an existing paid Fabric or Power BI Premium capacity, which independent estimates put at roughly $263/month minimum.
- The question is about your own data too, but you have no data platform to run it on, just a spreadsheet, a Postgres database or a warehouse and something you need charted, written up or presented → Julius AI. It is the only one of the eleven you can point at your own data for $0 and no platform contract: upload a file or connect Snowflake, BigQuery or Postgres, ask in plain language, and it returns charts, reports, presentations, websites, images or video out of one credits pool, with SOC 2 Type II certification and a Slack agent for asking from chat. Watch out for: no public API is documented on Julius's own site or in its docs, so it cannot be embedded in your own product; usage runs on an annual credits pool rather than flat unlimited access, so heavy analysis can hit a limit before the plan renews; and its own published pricing is unusually unstable: three independent write-ups quote the Plus plan at $16, $20 and $35 a month, with one reporting a recent $15 increase, so read julius.ai/pricing yourself rather than trusting any single review, this one included.
- You don't want a written research answer at all: you want structured data extracted from live web pages, or a browser agent that can complete a multi-step task on a page, callable by API or from Claude, Cursor or n8n → TinyFish. Search and Fetch return raw or LLM-ready markdown/JSON with no signup card required (150 URLs and 30 searches a minute, free indefinitely), while its Agent and Browser services complete multi-step web tasks and report an 85% pass rate and 99.3% detection coverage against anti-bot defenses. Watch out for: CAPTCHA handling on the toughest sites still leans on a third-party solver like CapSolver, independent pricing coverage disagrees with the vendor's own pay-as-you-go page, default concurrency is capped until you contact sales, and its headline accuracy numbers (98.7% accuracy, 89.9% Mind2Web score) come from the vendor or a single review, not an independently reproduced benchmark.
One pattern most confirm (with several mitigating it differently)
Eight of these eleven ship citations or a sourced answer, and four of them (Perplexity, GPT Researcher, Webhound and Elicit) carry the same underlying risk as a documented limitation rather than one an outside tester had to find: a cited answer is not a guaranteed-correct answer. Perplexity's own cons admit it "can still hallucinate or attach a citation that doesn't fully support the claim." Manus has been caught by independent testers producing fabricated "verification" output that looked real but described a state that didn't exist. Genspark's mitigation is architectural rather than a disclosure (cross-checking output across its nine blended models), but no independent audit was found confirming how well that holds up, and it stops short of claiming zero-hallucination reliability. GPT Researcher's own listing states plainly that "citations reduce but do not eliminate hallucination." Webhound's own blog is just as direct: it states plainly that building trust in AI-generated research remains an active, unsolved problem for the team, not a claim of solved reliability. Elicit's documentation goes furthest: it self-discloses hallucination as "a real risk despite safeguards" and directly instructs users to verify important findings against the original paper. Undermind's underlying LLM has drawn the same bias/ethics questions independent peer review raises for any GPT-class model, on top of its separately-documented inability to produce a reproducible search strategy. GC AI takes a different mitigation path rather than a disclosure: its research agents cross-check and reconcile each other's parallel findings before an answer is delivered, per its own documentation, a structural check most of the other six don't build in, though GC AI doesn't claim zero-hallucination reliability either. Microsoft Fabric Data Agent sits outside this pattern in an important way: it generates a query rather than a free-text answer, and that query runs against your own governed data rather than an external corpus. The honest risk there is a malformed or misdirected SQL/DAX/KQL query rather than a fabricated citation, and Microsoft's own documentation discloses the mitigations (read-only enforcement, per-user permission scoping, response caps) rather than a hallucination rate. Julius AI is outside it for the same structural reason and carries a different disclosure entirely: it analyses data you supply rather than a corpus it went and found, so its own listing's cons are about credit caps, an undocumented API and an unstable price list, not about fabricated sources. TinyFish sits outside the pattern for a related but distinct reason: it doesn't generate an answer or a citation at all, it extracts and returns raw web data or completes a browser action, so its own disclosed risk is a wrong extraction or a failed anti-bot bypass rather than a fabricated source, and its cons already flag that its headline accuracy figures (98.7% accuracy, 89.9% Mind2Web) come from the vendor or one review, not an independent reproduction. None of the eight is a substitute for checking the underlying data yourself; each is a different way of getting you there faster. The honest takeaway for a buyer isn't "which one is most accurate": it's which one gets you to a checkable source fastest for the job you actually have.
The pricing models split cleanly too, and it tracks the job split above rather than company size: GPT Researcher is the only one of the eleven with no subscription at all (open-source, pay only your own API usage); Perplexity, Elicit, Undermind and Julius AI are conventional freemium SaaS with flat monthly tiers, though Julius meters those tiers with an annual credits pool underneath and its published rates have moved between reviews; Manus and Genspark both run on usage-based credits instead, the same unpredictable-cost pattern our pricing census found concentrated in agents that execute open-ended multi-step work rather than answer bounded questions; Genspark's credits additionally expire unused at the end of each billing cycle, where Manus's do not carry the same disclosed expiry; Webhound is usage-based too, but priced directly in dollars against actual compute rather than an opaque credit block, with a one-time $5 signup credit rather than an ongoing free tier; TinyFish is also priced directly in dollars per step or minute rather than credits, but with an ongoing free tier instead of a signup credit: Search and Fetch stay free indefinitely, while Agent and Browser draw down a prepaid Wallet; GC AI has no free tier at all, just a 14-day trial ahead of a flat $500/month-and-up subscription, priced for an in-house team's budget rather than an individual's; and Microsoft Fabric Data Agent has no standalone price or free tier of any kind: it bills through whatever Fabric capacity you already own (or need to buy), an entirely different pricing shape from the other nine's per-seat, per-credit or per-dollar models.
FAQ
Which of these eleven is actually free?
GPT Researcher is the only one with no paid tier at all: it's open-source and free forever, though you pay your own LLM and search API usage. Perplexity, Manus, Genspark, Elicit, Undermind and Julius AI each have a free tier bolted onto a paid product, with Manus's free tier capped at roughly one task a day via its daily credit refill and Genspark's at about 100 credits a day. TinyFish splits down the middle: Search and Fetch remain free indefinitely with no card required (150 URLs, 30 searches a minute), while Agent and Browser draw against a paid, prepaid Wallet. Webhound, GC AI and Microsoft Fabric Data Agent have no ongoing free tier at all: Webhound gives new accounts a one-time $5 signup credit, GC AI offers only a 14-day trial on its $500/month-and-up plans, and Fabric Data Agent requires an existing paid Fabric or Power BI Premium capacity before you can use it at all.
Which one should an academic researcher use?
It depends which end of the review you're on. Elicit is purpose-built for screening a candidate set of scientific papers, with a 138M+-paper index, a Systematic Review workflow that screens thousands of papers against inclusion criteria, and structured data-extraction tables. Undermind is purpose-built for the earlier step (discovering papers you didn't know to look for via full-text reading and citation-trail following), but, per independent review, can't produce a documented search strategy the way Elicit's workflow can. Perplexity, Manus, Genspark, GPT Researcher and Webhound all research the open web, not specifically the academic literature; TinyFish also touches the open web but isn't a research tool in this sense at all: it's a data-extraction and browser-automation API, not a citation-generating answer engine; GC AI is scoped to legal rather than scientific research; and Microsoft Fabric Data Agent and Julius AI don't touch any public corpus at all, though Julius is the one worth knowing about for the step after the review, since it will run the analysis on a dataset you already hold and chart it.
What's the actual difference between Elicit and Undermind: don't they both search papers?
Yes, but for opposite jobs. Elicit takes inclusion criteria you define and screens a paper set against them at scale (up to 40,000 papers on Enterprise), producing a documented, auditable process and offering an API. Undermind instead tries to surface papers a keyword search would miss in the first place, reading full texts and following citation trails, and reports a "discovery curve" completeness estimate, but per an independent peer-reviewed review, it cannot produce a reproducible search strategy (its self-described "Achilles heel") and has no API. Use Elicit when you already know roughly what you're looking for and need rigor; use Undermind when you don't yet know what you're missing.
Which one can actually build and deploy something, not just write a report?
Manus is the only one of the eleven that writes, runs and deploys a small web app from an open-ended task, not just a compiled research report. Genspark produces slide decks and spreadsheets and can browse and build simple sites via its autopilot browser, but no source found documents it deploying a live, running web app the way Manus does. Julius AI also outputs more than a document (its own listing names presentations, websites, images and video), but those are built from the dataset you gave it, not from an arbitrary task you delegated.
Which one can I self-host?
GPT Researcher runs entirely on your own infrastructure with no vendor at all: free and open-source end to end. TinyFish's corpus record lists its deployment as both cloud and self-host, and its own published use cases describe an "open-source TinyFish agent" run locally, but its Search, Fetch and Browser services remain hosted APIs behind a paid Wallet, so self-hosting doesn't replace the vendor the way it does with GPT Researcher; whether the Agent component itself is openly licensed is not independently confirmed. Perplexity, Manus, Genspark, Webhound, Elicit, Undermind, GC AI, Microsoft Fabric Data Agent and Julius AI are all closed-source, cloud-hosted products with no self-host option at any tier: Fabric Data Agent specifically runs only inside a Fabric or Power BI Premium capacity, never on your own infrastructure.
Which one should an in-house legal team use?
GC AI: it's the only listing in this category scoped to US case law, statutes, regulations and agency guidance, running specialized agents in parallel across jurisdictions and courts before cross-checking and reconciling the findings. The other nine all research the open web, academic literature, general tasks or a company's own data, not legal-specific corpora. Note that on GC AI's cheapest $500/month plan, case-law research itself is a paid add-on rather than included by default: only the Team plan includes it out of the box.
Which one should I use to query my own company's data instead of the public internet?
Two of the eleven do this job and the choice between them is governance versus getting started. Microsoft Fabric Data Agent queries a Fabric lakehouse, warehouse or Power BI model you already run, scoping every query to the asking user's own permissions and Microsoft Purview policy, but it requires an existing paid Fabric or Power BI Premium capacity (independently estimated at roughly $263/month minimum) and cannot read unstructured documents like PDFs. Julius AI asks for no platform at all: upload a file or connect Snowflake, BigQuery or Postgres on a free account, and it returns charts, reports or presentations, with SOC 2 Type II certification behind it, the tradeoffs being an annual credits cap, no documented public API, and a price list independent write-ups have quoted three different ways.
Do any of these overlap enough to be a real head-to-head?
Not really: that's the point of this guide. The closest overlap is Manus versus Genspark, since both are broad, credit-billed autonomous agents that research and build deliverables rather than answer one bounded question; the real difference is surface and access: Manus ships a public developer API and browser-operator-driven app deployment, Genspark ships a wider consumer surface (slides, sheets, real phone calls, a personal-assistant layer) and no confirmed API. Perplexity's Deep Research mode also overlaps with GPT Researcher and Webhound, since all three plan sub-questions and compile a cited report from web search. The difference there is control, cost and delivery: Perplexity is a polished hosted product with a subscription, GPT Researcher is a free tool you self-host and point at whichever model and search engine you choose, and Webhound is a hosted, agent-triggered alternative priced per dollar of compute with a structured-Dataset output option neither of the other two offers. Elicit and Undermind overlap in domain (academic literature) but not in job (screening versus discovery), as covered above. The one genuine head-to-head added since this guide was first written is Microsoft Fabric Data Agent versus Julius AI: both answer questions about your own data rather than any public corpus, and they differ on entry rather than on capability: Fabric inherits the permissions and Purview policy of a platform you already pay for, while Julius connects to a warehouse or an uploaded spreadsheet from a free account and hands back a presentation. GC AI overlaps with neither, and with none of the other seven: it is the only one scoped to a body of law. TinyFish overlaps with none of the other ten either: it's the only one built to extract raw web data or automate a browser action rather than produce any kind of research answer.
Every agent above carries its own full researched listing: pricing tiers, sourced pros and cons, and a committed verdict on who it's for and who it isn't. Start with the research agents category for the full set, or see our State of AI Agents 2026 data report for how this category compares to the rest of the index.
Perplexity
Captured · theagentsindex.com Manus
Captured · theagentsindex.com Genspark
Captured · theagentsindex.com GPT Researcher
Captured · theagentsindex.com Webhound
Captured · theagentsindex.com Elicit
Captured · theagentsindex.com Undermind
Captured · theagentsindex.com GC AI
Captured · theagentsindex.com Julius AI
Captured · theagentsindex.com TinyFish
Captured · theagentsindex.com
Frequently asked questions
Which of these eleven is actually free?
Which one should an academic researcher use?
What's the actual difference between Elicit and Undermind: don't they both search papers?
Which one can actually build and deploy something, not just write a report?
Which one can I self-host?
Which one should an in-house legal team use?
Which one should I use to query my own company's data instead of the public internet?
Do any of these overlap enough to be a real head-to-head?
Advertise here
Reach buyers mid-decision. Reach builders choosing their next agent. Promote your brand with a display placement or bring your listing into focus with Featured.
Explore owner options →Advertise on this page →The digestFree
Which agents actually ship.
What we re-checked, what got added, and one number from the index. Tuesdays.
One-click unsubscribe