mirror of https://github.com/BerriAI/litellm.git synced 2025-04-26 03:04:13 +00:00

History

Krish Dholakia 56e9047818 Litellm router max depth (#6501 ) * feat(router.py): add check for max fallback depth Prevent infinite loop for fallbacks Closes https://github.com/BerriAI/litellm/issues/6498 * test: update test * (fix) Prometheus - Log Postgres DB latency, status on prometheus (#6484) * fix logging DB fails on prometheus * unit testing log to otel wrapper * unit testing for service logger + prometheus * use LATENCY buckets for service logging * fix service logging * docs clarify vertex vs gemini * (router_strategy/) ensure all async functions use async cache methods (#6489) * fix router strat * use async set / get cache in router_strategy * add coverage for router strategy * fix imports * fix batch_get_cache * use async methods for least busy * fix least busy use async methods * fix test_dual_cache_increment * test async_get_available_deployment when routing_strategy="least-busy" * (fix) proxy - fix when `STORE_MODEL_IN_DB` should be set (#6492) * set store_model_in_db at the top * correctly use store_model_in_db global * (fix) `PrometheusServicesLogger` `_get_metric` should return metric in Registry (#6486) * fix logging DB fails on prometheus * unit testing log to otel wrapper * unit testing for service logger + prometheus * use LATENCY buckets for service logging * fix service logging * fix _get_metric in prom services logger * add clear doc string * unit testing for prom service logger * bump: version 1.51.0 → 1.51.1 * Add `azure/gpt-4o-mini-2024-07-18` to model_prices_and_context_window.json (#6477) * Update utils.py (#6468) Fixed missing keys * (perf) Litellm redis router fix - ~100ms improvement (#6483) * docs(exception_mapping.md): add missing exception types Fixes https://github.com/Aider-AI/aider/issues/2120#issuecomment-2438971183 * fix(main.py): register custom model pricing with specific key Ensure custom model pricing is registered to the specific model+provider key combination * test: make testing more robust for custom pricing * fix(redis_cache.py): instrument otel logging for sync redis calls ensures complete coverage for all redis cache calls * refactor: pass parent_otel_span for redis caching calls in router allows for more observability into what calls are causing latency issues * test: update tests with new params * refactor: ensure e2e otel tracing for router * refactor(router.py): add more otel tracing acrosss router catch all latency issues for router requests * fix: fix linting error * fix(router.py): fix linting error * fix: fix test * test: fix tests * fix(dual_cache.py): pass ttl to redis cache * fix: fix param * perf(cooldown_cache.py): improve cooldown cache, to store cache results in memory for 5s, prevents redis call from being made on each request reduces 100ms latency per call with caching enabled on router * fix: fix test * fix(cooldown_cache.py): handle if a result is None * fix(cooldown_cache.py): add debug statements * refactor(dual_cache.py): move to using an in-memory check for batch get cache, to prevent redis from being hit for every call * fix(cooldown_cache.py): fix linting erropr * build: merge main --------- Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com> Co-authored-by: Xingyao Wang <xingyao@all-hands.dev> Co-authored-by: vibhanshu-ob <115142120+vibhanshu-ob@users.noreply.github.com>		2024-10-29 22:05:41 -07:00
..
_experimental	redis otel tracing + async support for latency routing (#6452 )	2024-10-28 21:52:12 -07:00
analytics_endpoints	Litellm ruff linting enforcement (#5992 )	2024-10-01 19:44:20 -04:00
auth	Litellm router max depth (#6501 )	2024-10-29 22:05:41 -07:00
common_utils	(code quality) add ruff check PLR0915 for `too-many-statements` (#6309 )	2024-10-18 15:36:49 +05:30
config_management_endpoints	feat(ui): for adding pass-through endpoints	2024-08-15 21:58:11 -07:00
db	(code quality) add ruff check PLR0915 for `too-many-statements` (#6309 )	2024-10-18 15:36:49 +05:30
example_config_yaml	Litellm router max depth (#6501 )	2024-10-29 22:05:41 -07:00
fine_tuning_endpoints	Add pyright to ci/cd + Fix remaining type-checking errors (#6082 )	2024-10-05 17:04:00 -04:00
guardrails	(code quality) add ruff check PLR0915 for `too-many-statements` (#6309 )	2024-10-18 15:36:49 +05:30
health_endpoints	(fix) Langfuse key based logging (#6372 )	2024-10-23 18:24:22 +05:30
hooks	redis otel tracing + async support for latency routing (#6452 )	2024-10-28 21:52:12 -07:00
management_endpoints	feat(custom_logger.py): expose new `async_dataset_hook` for modifying… (#6331 )	2024-10-20 09:00:04 -07:00
management_helpers	fix create_audit_log_for_update	2024-10-25 16:48:25 +04:00
openai_files_endpoints	Litellm ruff linting enforcement (#5992 )	2024-10-01 19:44:20 -04:00
pass_through_endpoints	(code quality) add ruff check PLR0915 for `too-many-statements` (#6309 )	2024-10-18 15:36:49 +05:30
proxy_load_test	Litellm ruff linting enforcement (#5992 )	2024-10-01 19:44:20 -04:00
rerank_endpoints	LiteLLM Minor Fixes & Improvements (09/26/2024) (#5925 ) (#5937 )	2024-09-27 17:54:13 -07:00
spend_tracking	(code quality) add ruff check PLR0915 for `too-many-statements` (#6309 )	2024-10-18 15:36:49 +05:30
ui_crud_endpoints	ui - add Create, get, delete endpoints for IP Addresses	2024-07-09 15:12:08 -07:00
vertex_ai_endpoints	feat(custom_logger.py): expose new `async_dataset_hook` for modifying… (#6331 )	2024-10-20 09:00:04 -07:00
.gitignore
__init__.py
_logging.py	fix(_logging.py): fix timestamp format for json logs	2024-06-20 15:20:21 -07:00
_new_secret_config.yaml	Litellm router max depth (#6501 )	2024-10-29 22:05:41 -07:00
_super_secret_config.yaml	docs(enterprise.md): cleanup docs	2024-07-15 14:52:08 -07:00
_types.py	Merge pull request #6433 from BerriAI/litellm_fix_audit_logs	2024-10-26 10:01:01 +04:00
cached_logo.jpg	(feat) use hosted images for custom branding	2024-02-22 14:51:40 -08:00
caching_routes.py	(refactor) caching use LLMCachingHandler for async_get_cache and set_cache (#6208 )	2024-10-14 16:34:01 +05:30
custom_sso.py	Litellm ruff linting enforcement (#5992 )	2024-10-01 19:44:20 -04:00
enterprise	feat(llama_guard.py): add llama guard support for content moderation + new `async_moderation_hook` endpoint	2024-02-17 19:13:04 -08:00
health_check.py	LiteLLM Minor Fixes and Improvements (09/14/2024) (#5697 )	2024-09-14 10:32:39 -07:00
lambda.py
litellm_pre_call_utils.py	LiteLLM Minor Fixes & Improvements (10/24/2024) (#6421 )	2024-10-25 15:55:56 -07:00
llamaguard_prompt.txt	feat(llama_guard.py): allow user to define custom unsafe content categories	2024-02-17 17:42:47 -08:00
logo.jpg	(feat) admin ui custom branding	2024-02-21 17:34:42 -08:00
openapi.json
post_call_rules.py
prisma_migration.py	Litellm expose disable schema update flag (#6085 )	2024-10-05 21:26:51 -04:00
proxy_cli.py	(docs + testing) Correctly document the timeout value used by litellm proxy is 6000 seconds + add to best practices for prod (#6339 )	2024-10-23 14:09:35 +05:30
proxy_config.yaml	(fix) `PrometheusServicesLogger` `_get_metric` should return metric in Registry (#6486 )	2024-10-29 21:29:19 +05:30
proxy_server.py	(fix) proxy - fix when `STORE_MODEL_IN_DB` should be set (#6492 )	2024-10-29 21:28:14 +05:30
README.md	[Feat-Proxy] Allow using custom sso handler (#5809 )	2024-09-20 19:14:33 -07:00
route_llm_request.py	(feat) use regex pattern matching for wildcard routing (#6150 )	2024-10-10 18:24:16 +05:30
schema.prisma	track created, updated at virtual keys	2024-10-25 07:19:29 +04:00
start.sh
utils.py	(fix) `PrometheusServicesLogger` `_get_metric` should return metric in Registry (#6486 )	2024-10-29 21:29:19 +05:30

README.md

litellm-proxy

A local, fast, and lightweight OpenAI-compatible server to call 100+ LLM APIs.

usage

$ pip install litellm

$ litellm --model ollama/codellama 

#INFO: Ollama running on http://0.0.0.0:8000

replace openai base

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

See how to call Huggingface,Bedrock,TogetherAI,Anthropic, etc.

Folder Structure

Routes

proxy_server.py - all openai-compatible routes - /v1/chat/completion, /v1/embedding + model info routes - /v1/models, /v1/model/info, /v1/model_group_info routes.
health_endpoints/ - /health, /health/liveliness, /health/readiness
management_endpoints/key_management_endpoints.py - all /key/* routes
management_endpoints/team_endpoints.py - all /team/* routes
management_endpoints/internal_user_endpoints.py - all /user/* routes
management_endpoints/ui_sso.py - all /sso/* routes