GPT Researcher does what the deep research modes in ChatGPT and Gemini do, in code you can run yourself. Give it a question and a planner agent breaks it into sub-questions, crawler agents search and scrape sources for each one in parallel, and a writer assembles a report, typically more than 2,000 words drawn from over 20 sources, with citations back to them. It can research the web, your local documents, or both, and export the result to PDF, Word and other formats.
A Deep Research mode explores a topic as a tree, following subtopics to a configurable depth and breadth; the README puts a run at about five minutes and $0.40 using OpenAI's o3-mini. It works as a web app with a lightweight or a Next.js front end, as a pip package for your own Python code, as an MCP client that can pull from sources such as GitHub, and as a Claude skill.
The project was started by Assaf Elovic in May 2023, is Apache-2.0 licensed, and has about 30,000 stars. It defaults to OpenAI for the model and Tavily for web search, and other model providers and search engines can be swapped in through configuration.
- Repository: github.com/assafelovic/gpt-researcher
- Licence: Apache-2.0 (Apache License 2.0)
- Language: Python. Stars: 29.9K. Forks: 4,080. Last push: Oct 1, 2026.
- Scan: safe, Sep 26, 2026, commit 0957c30
Who it is for
Analysts, students, journalists and developers who need a sourced overview of a topic quickly, and builders who want a research agent they can customize, point at internal documents, or embed in their own product.
Getting started
1. Clone the project (needs Python 3.12 or later)
git clone https://github.com/assafelovic/gpt-researcher.git && cd gpt-researcher2. Set your model and search API keys
export OPENAI_API_KEY=your_key && export TAVILY_API_KEY=your_key3. Install and start the server, then open http://localhost:8000
pip install -r requirements.txt && python -m uvicorn main:app --reload4. Or use it as a library in your own Python code
pip install gpt-researcherThe default setup needs two keys, OpenAI for the model and Tavily for web search; OPENAI_BASE_URL points it at another OpenAI-compatible or local model instead. Each report makes many model calls, so it costs more than a single chat. The README also has a Docker Compose setup.
Safety scan
We cloned assafelovic/gpt-researcher at commit 0957c30 on Sep 26, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- No secrets, no suspicious patterns, no install hooks and no committed binaries across 839 files and about 100,000 lines. The three bare-IP URLs are in tests/test_url_security.py, which checks that public addresses are allowed and private ones are blocked before the scraper fetches a page.
- There is no lockfile: requirements.txt and the package metadata use version ranges, so the scan had nothing exact to look up in OSV. pip resolves current versions when you install, which avoids stale pins but means two installs can differ.
- The hosts are what a research agent should touch: research sources such as arXiv, OpenAlex and news sites, model providers, and test domains like a.example. By default, queries go to OpenAI and Tavily; if you set TYPESAFE_API_KEY, scraped passages also go to TypeSafe's Jev service for filtering. We found no analytics.
- The scraper fetches pages it chooses from search results, so run the web server on localhost or behind authentication rather than on an open port.
- Six workflows, none using pull_request_target; only one of 10 third-party actions is pinned to a commit. Security policy, Dependabot, licence, contributing guide and code of conduct present; no CodeQL.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | None found. |
| Suspicious code | None found. |
| Install-time code | None: nothing runs at install beyond the package manager itself. |
| Committed binaries | None. |
| CI workflows | 6 workflows. None use pull_request_target. 9 of 10 third-party actions pinned to a tag rather than a commit. |
| Network hosts | 40 distinct hosts referenced from source; most often a.example, b.example, nature.com, arxiv.org. 3 URLs to a bare IP address, listed under the raw findings. |
| Known vulnerabilities | No lockfile to check: dependencies are declared as ranges, so what gets installed is whatever is current on the day. |
| Project hygiene | Has security policy, automated dependency updates, licence file, contributing guide. Missing CodeQL. |
| OpenSSF Scorecard | Not scored: the project is not in Scorecard's weekly index. |
The raw findings
Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.
URLs to bare IP addresses (3)
| Where | Rule | Match |
|---|---|---|
| tests/test_url_security.py:65 | ip-literal-url | assert validate_url("http://8.8.8.8/") == "http://8.8.8.8/" |
| tests/test_url_security.py:65 | ip-literal-url | assert validate_url("http://8.8.8.8/") == "http://8.8.8.8/" |
| tests/test_url_security.py:66 | ip-literal-url | assert is_safe_url("https://1.1.1.1/path?q=1") is True |
By the numbers
| Stars | 29.9K |
|---|---|
| Forks | 4,080 |
| Contributors | 285 |
| Commits | 3,211 |
| Open issues | 6 |
| Open pull requests | 17 |
| Releases | 74 |
| Latest release | v3.7.0 |
| Licence | Apache-2.0 |
| Main language | Python |
| Project age | 3 years |
| Last push | Oct 1, 2026 |
| Tracked files | 839 |
| Lines of code | 99.6K |
| Checkout size | 27 MB |
Lines by language: Python 37.4K, JSON 23.5K, Markdown 18.4K, TypeScript 10.2K, CSS 4,584, JavaScript 3,749.
Questions
Is GPT Researcher free?
The software is Apache-2.0 licensed and free. Running it costs whatever the model and search APIs charge; the README estimates about $0.40 for a Deep Research run on o3-mini. A local model behind an OpenAI-compatible endpoint removes the model cost.
Can GPT Researcher use my own documents?
Yes. It can research local files such as PDFs, Word documents, spreadsheets and Markdown alongside the web or instead of it, and through MCP it can pull from sources such as GitHub repositories and databases.
How accurate are the reports?
Statements are tied to sources and the citations make them checkable, but the agent still depends on what it finds and how the model summarizes it. Treat a report as a well-sourced first draft and follow the links for anything you plan to rely on.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
