LM/Performance

From Fundamental Ramen
< LM
Revision as of 09:17, 30 September 2026 by Tacoball (talk | contribs) (→Quick Test: llama-benchy)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Quick Test: llama-benchy

Use

# concurrency 8, test vLLM throughput
llama-benchy --base-url http://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --pp 512 --tg 128 --concurrency 8 \
  2>/dev/null | grep "^|" -B1 -A3

# concurrency 1
llama-benchy --base-url https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --pp 512 --tg 128

Install

pipx install llama-benchy

Full Test: betterbench

Use

betterbench run \
  --endpoint https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --out=path/to/bench-$(date +"%m%d")

Install

git clone https://github.com/GGZ14/BetterBench.git
uv sync
.venv/bin/activate