LM

From Fundamental Ramen
Jump to navigation Jump to search

Topics

🤖 Agent  🧠 Model  🚀 Engine  ⚙️ GPU & Driver  📊 Evaluation
ZooCode
axe
MCP
hf (Model Management)
Qwen Image 2.1
Bonsai 2 27B
Qwen 3.8 27B
Gemma 4
Devstral Small 2 24B
Experience
vLLM
llama-swap
ComfyUI
llama.cpp for SYCL
llama.cpp for CUDA
llama.cpp for OpenVINO
llama.cpp for TurboQuant
SGLang
Intel Arc Pro B70
Install driver
Install oneAPI (SYCL)
Install OpenVINO
.bashrc for B70
Experiment
NVIDIA RTX 3060
Install driver
AMD Ryzen AI 7 350
Experiment
Tools
Install nvtop
PCIe Trouble Shooting
Comparison
LM/GPU Comparison
Performance

News

Models management

File Path/images/quickmmd/LM-models-handling-infra.svg
File Size15.91 KB
Last modified17:54:52
Time of SVG loading1.60 second(s)
Time of SVG convertion0.00 second(s)
Cache DetectedNo
MD5 of incoming syntax45978182992b9851e9b971945432eebb
MD5 of cached syntax84bcc9c6aa8e79615db3a4e7791c22e5
SVG EngineKroki API
Kroki API URLhttp://kroki:8000
Mermaid syntax Referencehttps://mermaid.js.org/intro/syntax-reference.html
QuickMMD Version1.0.0 - About QuickMMD
flowchart LR
  A(LiteLLM)
  B(vLLM<br>Qwen 3.8 27B)
  C1(Intel Arc Pro B70)
  C2(Intel Arc Pro B70)

  A --> |Forward| B
  B --> |Tensor Parallelism| C1 & C2
  C1 --- |XCCL| C2
  • llama-swap can launch model, cannot manage users.
  • LiteLLM cannot launch model, can manage users.

Investigating llama.cpp on Intel Arc Pro B70

File Path/images/quickmmd/LM-tinkering-b70.svg
File Size40.88 KB
Last modified16:57:52
Time of SVG loading0.00 second(s)
Time of SVG convertion0.00 second(s)
Cache DetectedYes
SVG EngineKroki API
Kroki API URLhttp://kroki:8000
Mermaid syntax Referencehttps://mermaid.js.org/intro/syntax-reference.html
QuickMMD Version1.0.0 - About QuickMMD
flowchart TD
  subgraph Hardware
    A1(Intel Arc Pro B70):::node-okay
    A2(Asus ProArt B850 WiFi NEO):::node-okay
    A3(AMD Ryzen 5 9600X):::node-okay
  end

  subgraph OS
    B1(Install<br>Ubuntu 26.04):::node-okay
    A1 & A2 & A3 --> B1
  end

  subgraph Driver
    C1(Install PPA driver<br>26.27.39122.14):::node-okay
    C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
    C3(Install OpenVINO<br>2026.3.0<br>hard to install):::node-warn
    C1 --> C2 & C3
    B1 --> C1
  end

  subgraph Runtime
    D1[Build llama.cpp for<br>SYCL]:::node-okay
    D2[Build llama.cpp for<br>OpenVINO<br>not stable yet]:::node-warn
    %% D3[Build SGLang for<br>SYCL]:::node-todo
    %% D4[Build SGLang for<br>OpenVINO]:::node-todo
    C2 --> D1
    C3 --> D2
    %% C2 --> D3
    %% C3 --> D4
  end

  subgraph Tools
    T1(Install<br>pipx)
    T2(Install<br>hf)
    T3(Install<br>nvtop)
    T4(Install<br>tmux)
    T5(Install<br>lact)
    T6(Install<br>llama-swap)
    T1 --> T2
    B1 --> T1
    B1 --> T3
    B1 --> T4
    B1 --> T5
    T5 --> T6
  end

  subgraph Models
    M1(Qwen3.8-27B)
    M2(Devstral-Small-2-24B)
    T2 --> M1 & M2
  end

  subgraph config
    V1(Optimize<br>.bashrc)
    V2(Optimize<br>lmsw.yaml)
    C2 --> V1
    C3 --> V1
    T2 --> V1
    T4 ---> V1
    T6 ---> V2
  end

  subgraph Service
    Z(service<br>llama-swap \<br>  -config lmsw.yaml \<br>  -listen 0.0.0.0:9876)
    M1 --> Z
    D1 --> Z
    V2 --> Z
  end

  classDef node-okay fill:#efe,stroke:#393
  classDef node-warn fill:#fff0e0,stroke:#d63
  classDef node-fail fill:#fee,stroke:#d33
  classDef node-todo fill:#eee,stroke:#777