llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-12-03 09:53:45 +00:00

History

Botao Chen e3edca7739 feat: [new open benchmark] Math 500 (#1538 ) ## What does this PR do? Created a new math_500 open-benchmark based on OpenAI's [Let's Verify Step by Step](https://arxiv.org/abs/2305.20050) paper and hugging face's [HuggingFaceH4/MATH-500](https://huggingface.co/datasets/HuggingFaceH4/MATH-500) dataset. The challenge part of this benchmark is to parse the generated and expected answer and verify if they are same. For the parsing part, we refer to [Minerva: Solving Quantitative Reasoning Problems with Language Models](https://research.google/blog/minerva-solving-quantitative-reasoning-problems-with-language-models/). To simply the parse logic, as the next step, we plan to also refer to what [simple-eval](https://github.com/openai/simple-evals) is doing, using llm as judge to check if the generated answer matches the expected answer or not ## Test Plan on sever side, spin up a server with open-benchmark template `llama stack run llama_stack/templates/open-benchamrk/run.yaml` on client side, issue an open benchmark eval request `llama-stack-client --endpoint xxx eval run-benchmark "meta-reference-math-500" --model-id "meta-llama/Llama-3.3-70B-Instruct" --output-dir "/home/markchen1015/" --num-examples 20` and get ther aggregated eval results <img width="238" alt="Screenshot 2025-03-10 at 7 57 04 PM" src="https://github.com/user-attachments/assets/2c9da042-3b70-470e-a7c4-69f4cc24d1fb" /> check the generated answer and the related scoring and they make sense		2025-03-10 20:38:28 -07:00
..
bedrock	Fix precommit check after moving to ruff (#927 )	2025-02-02 06:46:45 -08:00
common	build: format codebase imports using ruff linter (#1028 )	2025-02-13 10:06:21 -08:00
datasetio	build: format codebase imports using ruff linter (#1028 )	2025-02-13 10:06:21 -08:00
inference	feat(logging): implement category-based logging (#1362 )	2025-03-07 11:34:30 -08:00
kvstore	chore: made inbuilt tools blocking calls into async non blocking calls (#1509 )	2025-03-09 16:59:24 -07:00
memory	fix(deps): move chardet and pypdf imports inline where used (#1434 )	2025-03-06 17:09:14 -08:00
scoring	feat: [new open benchmark] Math 500 (#1538 )	2025-03-10 20:38:28 -07:00
telemetry	fix: Agent telemetry inputs/outputs should be structured (#1302 )	2025-02-27 23:06:37 -08:00
__init__.py	API Updates (#73 )	2024-09-17 19:51:35 -07:00