LM: Difference between revisions
Jump to navigation
Jump to search
| (17 intermediate revisions by the same user not shown) | |||
| Line 2: | Line 2: | ||
{| class="wikitable" | {| class="wikitable" | ||
! 🤖 Agent || 🧠 Model || 🚀 Engine || ⚙️ GPU & Driver || Evaluation | ! 🤖 Agent || 🧠 Model || 🚀 Engine || ⚙️ GPU & Driver || 📊 Evaluation | ||
|- | |- | ||
| style="width:200px;vertical-align:top" | | | style="width:200px;vertical-align:top" | | ||
| Line 10: | Line 10: | ||
| style="width:200px;vertical-align:top" | | | style="width:200px;vertical-align:top" | | ||
: [[LM/hf|hf (Model Management)]] | : [[LM/hf|hf (Model Management)]] | ||
: [[LM/Qwen Image 2.1|Qwen Image 2.1]] | |||
: [[LM/Bonsai 2 27B|Bonsai 2 27B]] | |||
: [[LM/Qwen 3.8 27B|Qwen 3.8 27B]] | : [[LM/Qwen 3.8 27B|Qwen 3.8 27B]] | ||
: [[LM/Gemma 4|Gemma 4]] | : [[LM/Gemma 4|Gemma 4]] | ||
: [[LM/Devstral Small 2 24B|Devstral Small 2 24B]] | : [[LM/Devstral Small 2 24B|Devstral Small 2 24B]] | ||
: [[LM/Experience|Experience]] | : [[LM/Experience|Experience]] | ||
| style="width:200px;vertical-align:top" | | | style="width:200px;vertical-align:top" | | ||
: [[LM/vLLM|vLLM]] | |||
: [[LM/llama-swap|llama-swap]] | : [[LM/llama-swap|llama-swap]] | ||
: [[LM/ComfyUI|ComfyUI]] | : [[LM/ComfyUI|ComfyUI]] | ||
: [[LM/llama.cpp for SYCL|llama.cpp for SYCL]] | : [[LM/llama.cpp for SYCL|llama.cpp for SYCL]] | ||
: [[LM/llama.cpp for CUDA|llama.cpp for CUDA]] | : [[LM/llama.cpp for CUDA|llama.cpp for CUDA]] | ||
: [[LM/llama.cpp for OpenVINO|llama.cpp for OpenVINO]] | : [[LM/llama.cpp for OpenVINO|llama.cpp for OpenVINO]] | ||
: [[LM/llama.cpp for TurboQuant|llama.cpp for TurboQuant]] | : [[LM/llama.cpp for TurboQuant|llama.cpp for TurboQuant]] | ||
: [[LM/SGLang|SGLang]] | : [[LM/SGLang|SGLang]] | ||
| style="width:200px;vertical-align:top" | | | style="width:200px;vertical-align:top" | | ||
| Line 40: | Line 38: | ||
: Tools | : Tools | ||
:: [[LM/Install nvtop|Install nvtop]] | :: [[LM/Install nvtop|Install nvtop]] | ||
:: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]] | :: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]] | ||
: Comparison | |||
:: [[LM/GPU Comparison]] | |||
| style="width:200px;vertical-align:top" | | | style="width:200px;vertical-align:top" | | ||
: [[LM/Performance|Performance]] | : [[LM/Performance|Performance]] | ||
|} | |} | ||
| Line 53: | Line 50: | ||
* [https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.3/linux/ OpenVINO 2026.3 supports Ubuntu 26] - 2026-08-04 | * [https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.3/linux/ OpenVINO 2026.3 supports Ubuntu 26] - 2026-08-04 | ||
= | = Models management = | ||
<quickmmd name="models-handling-infra"> | <quickmmd name="models-handling-infra"> | ||
flowchart LR | flowchart LR | ||
A(LiteLLM) | A(LiteLLM) | ||
B( | B(vLLM<br>Qwen 3.8 27B) | ||
C1( | C1(Intel Arc Pro B70) | ||
C2( | C2(Intel Arc Pro B70) | ||
A --> B | A --> |Forward| B | ||
B --> C1 & C2 | B --> |Tensor Parallelism| C1 & C2 | ||
C1 --- |XCCL| C2 | |||
</quickmmd> | </quickmmd> | ||
= | * llama-swap can launch model, cannot manage users. | ||
* LiteLLM cannot launch model, can manage users. | |||
= Investigating llama.cpp on Intel Arc Pro B70 = | |||
<quickmmd name="tinkering-b70"> | <quickmmd name="tinkering-b70"> | ||
| Line 84: | Line 85: | ||
C1(Install PPA driver<br>26.27.39122.14):::node-okay | C1(Install PPA driver<br>26.27.39122.14):::node-okay | ||
C2(Install oneAPI<br>2026.1.1.20260724):::node-okay | C2(Install oneAPI<br>2026.1.1.20260724):::node-okay | ||
C1 --> C2 | |||
C1 --> C2 | |||
B1 --> C1 | B1 --> C1 | ||
end | end | ||
| Line 91: | Line 91: | ||
subgraph Runtime | subgraph Runtime | ||
D1[Build llama.cpp for<br>SYCL]:::node-okay | D1[Build llama.cpp for<br>SYCL]:::node-okay | ||
C2 --> D1 | C2 --> D1 | ||
end | end | ||
| Line 117: | Line 111: | ||
subgraph Models | subgraph Models | ||
M1(Qwen3.8-27B) | M1(Qwen3.8-27B) | ||
T2 --> M1 | |||
T2 --> M1 | |||
end | end | ||
| Line 125: | Line 118: | ||
V2(Optimize<br>lmsw.yaml) | V2(Optimize<br>lmsw.yaml) | ||
C2 --> V1 | C2 --> V1 | ||
T2 --> V1 | T2 --> V1 | ||
T4 ---> V1 | T4 ----> V1 | ||
T6 ---> V2 | T6 ---> V2 | ||
end | end | ||
Latest revision as of 03:12, 2 October 2026
Topics
| 🤖 Agent | 🧠 Model | 🚀 Engine | ⚙️ GPU & Driver | 📊 Evaluation |
|---|---|---|---|---|
|
News
Models management
- llama-swap can launch model, cannot manage users.
- LiteLLM cannot launch model, can manage users.
Investigating llama.cpp on Intel Arc Pro B70