4 min read

GPT Researcher: An Open Deep-Research Agent With Citations (GitHub, Scanned)

An open-source research agent that searches the web and your files and writes a cited report.

GPT Researcher logo
✅
Scan: safe. Nothing malicious. One thing to know: dependencies are not pinned, so the scan could not check them for known advisories, and every research run sends your question to the model and search providers you configure. Scanned Sep 26, 2026; the full report is below.

GPT Researcher does what the deep research modes in ChatGPT and Gemini do, in code you can run yourself. Give it a question and a planner agent breaks it into sub-questions, crawler agents search and scrape sources for each one in parallel, and a writer assembles a report, typically more than 2,000 words drawn from over 20 sources, with citations back to them. It can research the web, your local documents, or both, and export the result to PDF, Word and other formats.

A Deep Research mode explores a topic as a tree, following subtopics to a configurable depth and breadth; the README puts a run at about five minutes and $0.40 using OpenAI's o3-mini. It works as a web app with a lightweight or a Next.js front end, as a pip package for your own Python code, as an MCP client that can pull from sources such as GitHub, and as a Claude skill.

The project was started by Assaf Elovic in May 2023, is Apache-2.0 licensed, and has about 30,000 stars. It defaults to OpenAI for the model and Tavily for web search, and other model providers and search engines can be swapped in through configuration.

Who it is for

Analysts, students, journalists and developers who need a sourced overview of a topic quickly, and builders who want a research agent they can customize, point at internal documents, or embed in their own product.

Getting started

1. Clone the project (needs Python 3.12 or later)

git clone https://github.com/assafelovic/gpt-researcher.git && cd gpt-researcher

2. Set your model and search API keys

export OPENAI_API_KEY=your_key && export TAVILY_API_KEY=your_key

3. Install and start the server, then open http://localhost:8000

pip install -r requirements.txt && python -m uvicorn main:app --reload

4. Or use it as a library in your own Python code

pip install gpt-researcher

The default setup needs two keys, OpenAI for the model and Tavily for web search; OPENAI_BASE_URL points it at another OpenAI-compatible or local model instead. Each report makes many model calls, so it costs more than a single chat. The README also has a Docker Compose setup.

Safety scan

We cloned assafelovic/gpt-researcher at commit 0957c30 on Sep 26, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.

  • No secrets, no suspicious patterns, no install hooks and no committed binaries across 839 files and about 100,000 lines. The three bare-IP URLs are in tests/test_url_security.py, which checks that public addresses are allowed and private ones are blocked before the scraper fetches a page.
  • There is no lockfile: requirements.txt and the package metadata use version ranges, so the scan had nothing exact to look up in OSV. pip resolves current versions when you install, which avoids stale pins but means two installs can differ.
  • The hosts are what a research agent should touch: research sources such as arXiv, OpenAlex and news sites, model providers, and test domains like a.example. By default, queries go to OpenAI and Tavily; if you set TYPESAFE_API_KEY, scraped passages also go to TypeSafe's Jev service for filtering. We found no analytics.
  • The scraper fetches pages it chooses from search results, so run the web server on localhost or behind authentication rather than on an open port.
  • Six workflows, none using pull_request_target; only one of 10 third-party actions is pinned to a commit. Security policy, Dependabot, licence, contributing guide and code of conduct present; no CodeQL.

What the scanner counted

CheckResult
SecretsNone found.
Suspicious codeNone found.
Install-time codeNone: nothing runs at install beyond the package manager itself.
Committed binariesNone.
CI workflows6 workflows. None use pull_request_target. 9 of 10 third-party actions pinned to a tag rather than a commit.
Network hosts40 distinct hosts referenced from source; most often a.example, b.example, nature.com, arxiv.org. 3 URLs to a bare IP address, listed under the raw findings.
Known vulnerabilitiesNo lockfile to check: dependencies are declared as ranges, so what gets installed is whatever is current on the day.
Project hygieneHas security policy, automated dependency updates, licence file, contributing guide. Missing CodeQL.
OpenSSF ScorecardNot scored: the project is not in Scorecard's weekly index.

The raw findings

Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.

URLs to bare IP addresses (3)
WhereRuleMatch
tests/test_url_security.py:65ip-literal-urlassert validate_url("http://8.8.8.8/") == "http://8.8.8.8/"
tests/test_url_security.py:65ip-literal-urlassert validate_url("http://8.8.8.8/") == "http://8.8.8.8/"
tests/test_url_security.py:66ip-literal-urlassert is_safe_url("https://1.1.1.1/path?q=1") is True

By the numbers

Stars29.9K
Forks4,080
Contributors285
Commits3,211
Open issues6
Open pull requests17
Releases74
Latest releasev3.7.0
LicenceApache-2.0
Main languagePython
Project age3 years
Last pushOct 1, 2026
Tracked files839
Lines of code99.6K
Checkout size27 MB

Lines by language: Python 37.4K, JSON 23.5K, Markdown 18.4K, TypeScript 10.2K, CSS 4,584, JavaScript 3,749.

Questions

Is GPT Researcher free?

The software is Apache-2.0 licensed and free. Running it costs whatever the model and search APIs charge; the README estimates about $0.40 for a Deep Research run on o3-mini. A local model behind an OpenAI-compatible endpoint removes the model cost.

Can GPT Researcher use my own documents?

Yes. It can research local files such as PDFs, Word documents, spreadsheets and Markdown alongside the web or instead of it, and through MCP it can pull from sources such as GitHub repositories and databases.

How accurate are the reports?

Statements are tied to sources and the citations make them checkable, but the agent still depends on what it finds and how the model summarizes it. Treat a report as a well-sourced first draft and follow the links for anything you plan to rely on.


This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.