feat: consolidate most distros into "starter"

* Removes a bunch of distros * Removed distros were added into the "starter" distribution * Doc for "starter" has been added * Partially reverts https://github.com/meta-llama/llama-stack/pull/2482 since inference providers are disabled by default and can be turned on manually via env variable. * Disables safety in starter distro Closes: #2502 Signed-off-by: Sébastien Han <seb@redhat.com>
2025-06-29 11:24:19 +00:00 · 2025-06-25 16:09:41 +02:00 · 2025-06-25 16:09:41 +02:00 · bedfea38c3
commit bedfea38c3
parent 0ddb293d77
127 changed files with 758 additions and 10771 deletions
--- a/llama_stack/distribution/providers.py
+++ b/llama_stack/distribution/providers.py
@ -84,7 +84,13 @@ class ProviderImpl(Providers):
                Each API maps to a dictionary of provider IDs to their health responses.
        """
        providers_health: dict[str, dict[str, HealthResponse]] = {}
-        timeout = 1.0
+
+        # The timeout has to be long enough to allow all the providers to be checked, especially in
+        # the case of the inference router health check since it checks all registered inference
+        # providers.
+        # The timeout must not be equal to the one set by health method for a given implementation,
+        # otherwise we will miss some providers.
+        timeout = 3.0

        async def check_provider_health(impl: Any) -> tuple[str, HealthResponse] | None:
            # Skip special implementations (inspect/providers) that don't have provider specs