llama-stack

forked from phoenix-oss/llama-stack-mirror

History

Dinesh Yeduguru 99bbe0e70b feat: Add new compact MetricInResponse type (#1593 ) # What does this PR do? This change adds a compact type to include metrics in response as opposed to the full MetricEvent which is relevant for internal logging purposes. ## Test Plan ``` LLAMA_STACK_CONFIG=~/.llama/distributions/fireworks/fireworks-run.yaml pytest -s -v agents/test_agents.py --safety-shield meta-llama/Llama-Guard-3-8B --text-model meta-llama/Llama-3.1-8B-Instruct llama stack run ~/.llama/distributions/fireworks/fireworks-run.yaml curl --request POST \ --url http://localhost:8321/v1/inference/chat-completion \ --header 'content-type: application/json' \ --data '{ "model_id": "meta-llama/Llama-3.1-70B-Instruct", "messages": [ { "role": "user", "content": { "type": "text", "text": "where do humans live" } } ], "stream": false }' { "metrics": [ { "metric": "prompt_tokens", "value": 10, "unit": null }, { "metric": "completion_tokens", "value": 522, "unit": null }, { "metric": "total_tokens", "value": 532, "unit": null } ], "completion_message": { "role": "assistant", "content": "Humans live in various parts of the world...............", "stop_reason": "out_of_tokens", "tool_calls": [] }, "logprobs": null } ```		2025-03-12 15:45:44 -07:00
..
routers	feat: Add new compact MetricInResponse type (#1593 )	2025-03-12 15:45:44 -07:00
server	feat: Add back inference metrics and preserve context variables across asyncio boundary (#1552 )	2025-03-12 12:01:03 -07:00
store	refactor: move a few tests to top-level tests/ directory	2025-03-03 17:33:39 -08:00
ui	docs: update test_agents to use new Agent SDK API (#1402 )	2025-03-06 15:21:12 -08:00
utils	fix: fix build error in context.py (#1595 )	2025-03-12 13:26:23 -07:00
__init__.py	API Updates (#73 )	2024-09-17 19:51:35 -07:00
build.py	refactor: `ImageType` to `LlamaStackImageType` (#1500 )	2025-03-10 17:12:53 -04:00
build_conda_env.sh	chore: remove straggler references to llama-models (#1345 )	2025-03-01 14:26:03 -08:00
build_container.sh	chore: remove straggler references to llama-models (#1345 )	2025-03-01 14:26:03 -08:00
build_venv.sh	chore: remove straggler references to llama-models (#1345 )	2025-03-01 14:26:03 -08:00
client.py	chore: move all Llama Stack types from llama-models to llama-stack (#1098 )	2025-02-14 09:10:59 -08:00
common.sh	fix: Fixing some small issues with the build scripts (#1132 )	2025-02-19 22:20:49 -08:00
configure.py	fix: resolve pydantic warning on .dict() usage (#1445 )	2025-03-06 11:27:47 -08:00
datatypes.py	fix!: update eval-tasks -> benchmarks (#1032 )	2025-02-13 16:40:58 -08:00
distribution.py	chore(lint): update Ruff ignores for project conventions and maintainability (#1184 )	2025-02-28 09:36:49 -08:00
inspect.py	fix: improve signal handling and update dependencies (#1044 )	2025-02-13 08:07:59 -08:00
library_client.py	feat: Add back inference metrics and preserve context variables across asyncio boundary (#1552 )	2025-03-12 12:01:03 -07:00
request_headers.py	feat: Add back inference metrics and preserve context variables across asyncio boundary (#1552 )	2025-03-12 12:01:03 -07:00
resolver.py	feat: Add back inference metrics and preserve context variables across asyncio boundary (#1552 )	2025-03-12 12:01:03 -07:00
stack.py	feat(logging): implement category-based logging (#1362 )	2025-03-07 11:34:30 -08:00
start_stack.sh	feat(logging): implement category-based logging (#1362 )	2025-03-07 11:34:30 -08:00