LM/Performance
< LM
Quick Test: llama-benchy
Use
# concurrency 8, test vLLM throughput
llama-benchy --base-url http://models.nothanks.com/v1 \
--model my_model --api-key sk-12345678 \
--pp 512 --tg 128 --concurrency 8 \
2>/dev/null | grep "^|" -B1 -A3
# concurrency 1
llama-benchy --base-url https://models.nothanks.com/v1 \
--model my_model --api-key sk-12345678 \
--pp 512 --tg 128
Install
pipx install llama-benchy
Full Test: betterbench
Use
betterbench run \
--endpoint https://models.nothanks.com/v1 \
--model my_model --api-key sk-12345678 \
--out=path/to/bench-$(date +"%m%d")
Install
git clone https://github.com/GGZ14/BetterBench.git
uv sync
.venv/bin/activate