phoenix-oss/llama-stack-mirror

Fork 1

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-07-17 10:28:11 +00:00

Kelly Brown b096794959

SqlStore Integration Tests / test-postgres (3.13) (push) Failing after 2s

Details

Integration Tests / discover-tests (push) Successful in 2s

Details

Vector IO Integration Tests / test-matrix (3.12, inline::milvus) (push) Failing after 17s

Details

Integration Auth Tests / test-matrix (oauth2_token) (push) Failing after 19s

Details

Python Package Build Test / build (3.12) (push) Failing after 14s

Details

Test Llama Stack Build / build-custom-container-distribution (push) Failing after 14s

Details

Vector IO Integration Tests / test-matrix (3.12, remote::pgvector) (push) Failing after 15s

Details

SqlStore Integration Tests / test-postgres (3.12) (push) Failing after 20s

Details

Unit Tests / unit-tests (3.13) (push) Failing after 15s

Details

Test Llama Stack Build / generate-matrix (push) Successful in 16s

Details

Vector IO Integration Tests / test-matrix (3.13, remote::pgvector) (push) Failing after 20s

Details

Test External Providers / test-external-providers (venv) (push) Failing after 17s

Details

Update ReadTheDocs / update-readthedocs (push) Failing after 15s

Details

Test Llama Stack Build / build-single-provider (push) Failing after 21s

Details

Test Llama Stack Build / build-ubi9-container-distribution (push) Failing after 18s

Details

Unit Tests / unit-tests (3.12) (push) Failing after 22s

Details

Vector IO Integration Tests / test-matrix (3.12, inline::sqlite-vec) (push) Failing after 25s

Details

Vector IO Integration Tests / test-matrix (3.13, remote::chromadb) (push) Failing after 23s

Details

Vector IO Integration Tests / test-matrix (3.13, inline::milvus) (push) Failing after 26s

Details

Vector IO Integration Tests / test-matrix (3.13, inline::sqlite-vec) (push) Failing after 19s

Details

Vector IO Integration Tests / test-matrix (3.12, inline::faiss) (push) Failing after 28s

Details

Vector IO Integration Tests / test-matrix (3.13, inline::faiss) (push) Failing after 21s

Details

Vector IO Integration Tests / test-matrix (3.12, remote::chromadb) (push) Failing after 23s

Details

Python Package Build Test / build (3.13) (push) Failing after 44s

Details

Test Llama Stack Build / build (push) Failing after 25s

Details

Integration Tests / test-matrix (push) Failing after 46s

Details

Pre-commit / pre-commit (push) Successful in 2m24s

Details

docs: Reorganize documentation on the webpage (#2651 )

# What does this PR do?
Reorganizes the Llama stack webpage into more concise index pages,
introduce more of a workflow, and reduce repetition of content.

New nav structure so far based on #2637 

Further discussions in
https://github.com/meta-llama/llama-stack/discussions/2585

**Preview:**
![Screenshot 2025-07-09 at 2 31
53 PM](https://github.com/user-attachments/assets/4c1f3845-b328-4f12-9f20-3f09375007af)

You can also build a full local preview locally 

 **Feedback**
Looking for feedback on page titles and general feedback on the new
structure

**Follow up documentation**
I plan on reducing some sections and standardizing some terminology in a
follow up PR.
More discussions on that in
https://github.com/meta-llama/llama-stack/discussions/2585

2025-07-15 14:19:35 -07:00

4.9 KiB

Raw Blame History

Llama Stack

Welcome to Llama Stack, the open-source framework for building generative AI applications.

:class: tip

Check out [Getting Started with Llama 4](https://colab.research.google.com/github/meta-llama/llama-stack/blob/main/docs/getting_started_llama4.ipynb)

:class: tip

Llama Stack {{ llama_stack_version }} is now available! See the {{ llama_stack_version_link }} for more details.

What is Llama Stack?

Llama Stack defines and standardizes the core building blocks needed to bring generative AI applications to market. It provides a unified set of APIs with implementations from leading service providers, enabling seamless transitions between development and production environments. More specifically, it provides

Unified API layer for Inference, RAG, Agents, Tools, Safety, Evals, and Telemetry.
Plugin architecture to support the rich ecosystem of implementations of the different APIs in different environments like local development, on-premises, cloud, and mobile.
Prepackaged verified distributions which offer a one-stop solution for developers to get started quickly and reliably in any environment
Multiple developer interfaces like CLI and SDKs for Python, Node, iOS, and Android
Standalone applications as examples for how to build production-grade AI applications with Llama Stack

:alt: Llama Stack
:width: 400px

Our goal is to provide pre-packaged implementations (aka "distributions") which can be run in a variety of deployment environments. LlamaStack can assist you in your entire app development lifecycle - start iterating on local, mobile or desktop and seamlessly transition to on-prem or public cloud deployments. At every point in this transition, the same set of APIs and the same developer experience is available.

How does Llama Stack work?

Llama Stack consists of a server (with multiple pluggable API providers) and Client SDKs (see below) meant to be used in your applications. The server can be run in a variety of environments, including local (inline) development, on-premises, and cloud. The client SDKs are available for Python, Swift, Node, and Kotlin.

Quick Links

Ready to build? Check out the Quick Start to get started.
Want to contribute? See the Contributing guide.

Supported Llama Stack Implementations

A number of "adapters" are available for some popular Inference and Vector Store providers. For other APIs (particularly Safety and Agents), we provide reference implementations you can use to get started. We expect this list to grow over time. We are slowly onboarding more providers to the ecosystem as we get more confidence in the APIs.

Inference API

Provider	Environments
Meta Reference	Single Node
Ollama	Single Node
Fireworks	Hosted
Together	Hosted
NVIDIA NIM	Hosted and Single Node
vLLM	Hosted and Single Node
TGI	Hosted and Single Node
AWS Bedrock	Hosted
Cerebras	Hosted
Groq	Hosted
SambaNova	Hosted
PyTorch ExecuTorch	On-device iOS, Android
OpenAI	Hosted
Anthropic	Hosted
Gemini	Hosted
WatsonX	Hosted

Agents API

Provider	Environments
Meta Reference	Single Node
Fireworks	Hosted
Together	Hosted
PyTorch ExecuTorch	On-device iOS

Vector IO API

Provider	Environments
FAISS	Single Node
SQLite-Vec	Single Node
Chroma	Hosted and Single Node
Milvus	Hosted and Single Node
Postgres (PGVector)	Hosted and Single Node
Weaviate	Hosted
Qdrant	Hosted and Single Node

Safety API

Provider	Environments
Llama Guard	Depends on Inference Provider
Prompt Guard	Single Node
Code Scanner	Single Node
AWS Bedrock	Hosted

Post Training API

Provider	Environments
Meta Reference	Single Node
HuggingFace	Single Node
TorchTune	Single Node
NVIDIA NEMO	Hosted

Eval API

Provider	Environments
Meta Reference	Single Node
NVIDIA NEMO	Hosted

Telemetry API

Provider	Environments
Meta Reference	Single Node

Tool Runtime API

Provider	Environments
Brave Search	Hosted
RAG Runtime	Single Node

:hidden:
:maxdepth: 3

self
getting_started/index
concepts/index
providers/index
distributions/index
advanced_apis/index
building_applications/index
deploying/index
contributing/index
references/index

4.9 KiB Raw Blame History

Llama Stack

What is Llama Stack?

How does Llama Stack work?

Quick Links

Supported Llama Stack Implementations

4.9 KiB

Raw Blame History