forked from phoenix-oss/llama-stack-mirror
* wip scoring refactor * llm as judge, move folders * test full generation + eval * extract score regex to llm context * remove prints, cleanup braintrust in this branch * change json -> class * remove initialize * address nits * check identifier prefix * udpate MANIFEST |
||
|---|---|---|
| .. | ||
| agents | ||
| datasetio | ||
| eval | ||
| inference | ||
| memory | ||
| safety | ||
| scoring | ||
| __init__.py | ||
| resolver.py | ||