llama-stack

forked from phoenix-oss/llama-stack-mirror

History

Botao Chen 4dccf916d1 feat: open benchmark template and doc (#1465 ) ## What does this PR do? - Provide a distro template to let developer easily run the open benchmarks llama stack supports on llama and non-llama models. - Provide doc on how to run open benchmark eval via CLI and open benchmark contributing guide [//]: # (If resolving an issue, uncomment and update the line below) (Closes #1375 ) ## Test Plan open benchmark eval results on llama, gpt, gemini and clause <img width="771" alt="Screenshot 2025-03-06 at 7 33 05 PM" src="https://github.com/user-attachments/assets/1bd85456-b9b9-4b37-af76-4ce1d2bac00e" /> doc preview <img width="944" alt="Screenshot 2025-03-06 at 7 33 58 PM" src="https://github.com/user-attachments/assets/f4e5866d-b395-4c40-aa8b-080edeb5cdb6" /> <img width="955" alt="Screenshot 2025-03-06 at 7 34 04 PM" src="https://github.com/user-attachments/assets/629defb6-d5e4-473c-aa03-308bce386fb4" /> <img width="965" alt="Screenshot 2025-03-06 at 7 35 29 PM" src="https://github.com/user-attachments/assets/c21ff96c-9e8c-4c54-b6b8-25883125f4cf" /> <img width="957" alt="Screenshot 2025-03-06 at 7 35 37 PM" src="https://github.com/user-attachments/assets/47571c90-1381-4e2c-bbed-c4f3a60578d0" />		2025-03-07 10:37:55 -08:00
..
apis	fix: Revert "feat: record token usage for inference API (#1300 )" (#1476 )	2025-03-07 10:16:47 -08:00
cli	fix: resolve pydantic warning on .dict() usage (#1445 )	2025-03-06 11:27:47 -08:00
distribution	fix: Revert "feat: record token usage for inference API (#1300 )" (#1476 )	2025-03-07 10:16:47 -08:00
models/llama	refactor: move a few tests to top-level tests/ directory	2025-03-03 17:33:39 -08:00
providers	fix: Revert "feat: record token usage for inference API (#1300 )" (#1476 )	2025-03-07 10:16:47 -08:00
scripts	refactor(test): introduce --stack-config and simplify options (#1404 )	2025-03-05 17:02:02 -08:00
strong_typing	Ensure that deprecations for fields follow through to OpenAPI	2025-02-19 13:54:04 -08:00
templates	feat: open benchmark template and doc (#1465 )	2025-03-07 10:37:55 -08:00
__init__.py	export LibraryClient	2024-12-13 12:08:00 -08:00
env.py	refactor(test): move tools, evals, datasetio, scoring and post training tests (#1401 )	2025-03-04 14:53:47 -08:00
logcat.py	feat: add a configurable category-based logger (#1352 )	2025-03-02 18:51:14 -08:00
schema_utils.py	ci: add mypy for static type checking (#1101 )	2025-02-21 13:15:40 -08:00