LM: Difference between revisions
Jump to navigation
Jump to search
| Line 67: | Line 67: | ||
* LiteLLM cannot launch model, can manage users. | * LiteLLM cannot launch model, can manage users. | ||
= | = Investigating llama.cpp on Intel Arc Pro B70 = | ||
<quickmmd name="tinkering-b70"> | <quickmmd name="tinkering-b70"> | ||
Revision as of 03:08, 2 October 2026
Topics
| 🤖 Agent | 🧠 Model | 🚀 Engine | ⚙️ GPU & Driver | 📊 Evaluation |
|---|---|---|---|---|
|
News
Models management
| File Path | /images/quickmmd/LM-models-handling-infra.svg |
|---|---|
| File Size | 15.91 KB |
| Last modified | 17:54:52 |
| Time of SVG loading | 0.00 second(s) |
| Time of SVG convertion | 0.00 second(s) |
| Cache Detected | Yes |
| SVG Engine | Kroki API |
| Kroki API URL | http://kroki:8000 |
| Mermaid syntax Reference | https://mermaid.js.org/intro/syntax-reference.html |
| QuickMMD Version | 1.0.0 - About QuickMMD |
flowchart LR
A(LiteLLM)
B(vLLM<br>Qwen 3.8 27B)
C1(Intel Arc Pro B70)
C2(Intel Arc Pro B70)
A --> |Forward| B
B --> |Tensor Parallelism| C1 & C2
C1 --- |XCCL| C2
- llama-swap can launch model, cannot manage users.
- LiteLLM cannot launch model, can manage users.
Investigating llama.cpp on Intel Arc Pro B70
| File Path | /images/quickmmd/LM-tinkering-b70.svg |
|---|---|
| File Size | 40.88 KB |
| Last modified | 16:57:52 |
| Time of SVG loading | 0.00 second(s) |
| Time of SVG convertion | 0.00 second(s) |
| Cache Detected | Yes |
| SVG Engine | Kroki API |
| Kroki API URL | http://kroki:8000 |
| Mermaid syntax Reference | https://mermaid.js.org/intro/syntax-reference.html |
| QuickMMD Version | 1.0.0 - About QuickMMD |
flowchart TD
subgraph Hardware
A1(Intel Arc Pro B70):::node-okay
A2(Asus ProArt B850 WiFi NEO):::node-okay
A3(AMD Ryzen 5 9600X):::node-okay
end
subgraph OS
B1(Install<br>Ubuntu 26.04):::node-okay
A1 & A2 & A3 --> B1
end
subgraph Driver
C1(Install PPA driver<br>26.27.39122.14):::node-okay
C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
C3(Install OpenVINO<br>2026.3.0<br>hard to install):::node-warn
C1 --> C2 & C3
B1 --> C1
end
subgraph Runtime
D1[Build llama.cpp for<br>SYCL]:::node-okay
D2[Build llama.cpp for<br>OpenVINO<br>not stable yet]:::node-warn
%% D3[Build SGLang for<br>SYCL]:::node-todo
%% D4[Build SGLang for<br>OpenVINO]:::node-todo
C2 --> D1
C3 --> D2
%% C2 --> D3
%% C3 --> D4
end
subgraph Tools
T1(Install<br>pipx)
T2(Install<br>hf)
T3(Install<br>nvtop)
T4(Install<br>tmux)
T5(Install<br>lact)
T6(Install<br>llama-swap)
T1 --> T2
B1 --> T1
B1 --> T3
B1 --> T4
B1 --> T5
T5 --> T6
end
subgraph Models
M1(Qwen3.8-27B)
M2(Devstral-Small-2-24B)
T2 --> M1 & M2
end
subgraph config
V1(Optimize<br>.bashrc)
V2(Optimize<br>lmsw.yaml)
C2 --> V1
C3 --> V1
T2 --> V1
T4 ---> V1
T6 ---> V2
end
subgraph Service
Z(service<br>llama-swap \<br> -config lmsw.yaml \<br> -listen 0.0.0.0:9876)
M1 --> Z
D1 --> Z
V2 --> Z
end
classDef node-okay fill:#efe,stroke:#393
classDef node-warn fill:#fff0e0,stroke:#d63
classDef node-fail fill:#fee,stroke:#d33
classDef node-todo fill:#eee,stroke:#777