LM/Performance: Difference between revisions

From Fundamental Ramen
< LM
Jump to navigation Jump to search
 
(3 intermediate revisions by the same user not shown)
Line 4: Line 4:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
llama-benchy --base-url https://models.nothanks.com/v1 --pp 512 --tg 128 --model my_model --api-key sk-12345678
# concurrency 8, test vLLM throughput
 
llama-benchy --base-url http://models.nothanks.com/v1 \
llama-benchy --base-url http://models.nothanks.com/v1 \
   --model my_model --api-key sk-12345678 \
   --model my_model --api-key sk-12345678 \
   --pp 512 --tg 128 --concurrency 8 \
   --pp 512 --tg 128 --concurrency 8 \
   2>/dev/null | grep "^|" -B1 -A3
   2>/dev/null | grep "^|" -B1 -A3
# concurrency 1
llama-benchy --base-url https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --pp 512 --tg 128
</syntaxhighlight>
</syntaxhighlight>


Line 23: Line 27:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
betterbench run --endpoint https://models.nothanks.com/v1 --model my_model --api-key sk-12345678
betterbench run \
  --endpoint https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --out=path/to/bench-$(date +"%m%d")
</syntaxhighlight>
</syntaxhighlight>



Latest revision as of 09:17, 30 September 2026

Quick Test: llama-benchy

Use

# concurrency 8, test vLLM throughput
llama-benchy --base-url http://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --pp 512 --tg 128 --concurrency 8 \
  2>/dev/null | grep "^|" -B1 -A3

# concurrency 1
llama-benchy --base-url https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --pp 512 --tg 128

Install

pipx install llama-benchy

Full Test: betterbench

Use

betterbench run \
  --endpoint https://models.nothanks.com/v1 \
  --model my_model --api-key sk-12345678 \
  --out=path/to/bench-$(date +"%m%d")

Install

git clone https://github.com/GGZ14/BetterBench.git
uv sync
.venv/bin/activate