LM: Difference between revisions

From Fundamental Ramen
Jump to navigation Jump to search
 
(12 intermediate revisions by the same user not shown)
Line 2: Line 2:


{| class="wikitable"
{| class="wikitable"
! 🤖 Agent  || 🧠 Model  || 🚀 Engine  || ⚙️ GPU & Driver || Evaluation
! 🤖 Agent  || 🧠 Model  || 🚀 Engine  || ⚙️ GPU & Driver || 📊 Evaluation
|-
|-
| style="width:200px;vertical-align:top" |
| style="width:200px;vertical-align:top" |
Line 10: Line 10:
| style="width:200px;vertical-align:top" |
| style="width:200px;vertical-align:top" |
: [[LM/hf|hf (Model Management)]]
: [[LM/hf|hf (Model Management)]]
: [[LM/Qwen Image 2.1|Qwen Image 2.1]]
: [[LM/Bonsai 2 27B|Bonsai 2 27B]]
: [[LM/Qwen 3.8 27B|Qwen 3.8 27B]]
: [[LM/Qwen 3.8 27B|Qwen 3.8 27B]]
: [[LM/Gemma 4|Gemma 4]]
: [[LM/Gemma 4|Gemma 4]]
Line 15: Line 17:
: [[LM/Experience|Experience]]
: [[LM/Experience|Experience]]
| style="width:200px;vertical-align:top" |
| style="width:200px;vertical-align:top" |
: [[LM/vLLM|vLLM]]
: [[LM/llama-swap|llama-swap]]
: [[LM/llama-swap|llama-swap]]
: [[LM/vLLM|vLLM]]
: [[LM/ComfyUI|ComfyUI]]
: [[LM/ComfyUI|ComfyUI]]
: [[LM/llama.cpp for SYCL|llama.cpp for SYCL]]
: [[LM/llama.cpp for SYCL|llama.cpp for SYCL]]
Line 36: Line 38:
: Tools
: Tools
:: [[LM/Install nvtop|Install nvtop]]
:: [[LM/Install nvtop|Install nvtop]]
:: [[LM/lact|lact]]
:: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]]
:: [[LM/PCIe Trouble Shooting|PCIe Trouble Shooting]]
: Comparison
:: [[LM/GPU Comparison]]
| style="width:200px;vertical-align:top" |
| style="width:200px;vertical-align:top" |
: [[LM/Performance|Performance]]
: [[LM/Performance|Performance]]
: [[LM/Security|Security]]
: [[LM/Hallucination|Hallucination]]
|}
|}


Line 54: Line 55:
flowchart LR
flowchart LR
   A(LiteLLM)
   A(LiteLLM)
   B(llama-swap)
   B(vLLM<br>Qwen 3.8 27B)
   C1(llama.cpp)
   C1(Intel Arc Pro B70)
   C2(vLLM)
   C2(Intel Arc Pro B70)


   A --> B
   A --> |Forward| B
   B --> C1 & C2
   B --> |Tensor Parallelism| C1 & C2
  C1 --- |XCCL| C2
</quickmmd>
</quickmmd>


Line 65: Line 67:
* LiteLLM cannot launch model, can manage users.
* LiteLLM cannot launch model, can manage users.


= Decision Tree for tinkering with Intel Arc Pro B70 =
= Investigating llama.cpp on Intel Arc Pro B70 =


<quickmmd name="tinkering-b70">
<quickmmd name="tinkering-b70">
Line 83: Line 85:
     C1(Install PPA driver<br>26.27.39122.14):::node-okay
     C1(Install PPA driver<br>26.27.39122.14):::node-okay
     C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
     C2(Install oneAPI<br>2026.1.1.20260724):::node-okay
    C3(Install OpenVINO<br>2026.3.0<br>hard to install):::node-warn
     C1 --> C2
     C1 --> C2 & C3
     B1 --> C1
     B1 --> C1
   end
   end
Line 90: Line 91:
   subgraph Runtime
   subgraph Runtime
     D1[Build llama.cpp for<br>SYCL]:::node-okay
     D1[Build llama.cpp for<br>SYCL]:::node-okay
    D2[Build llama.cpp for<br>OpenVINO<br>not stable yet]:::node-warn
    %% D3[Build SGLang for<br>SYCL]:::node-todo
    %% D4[Build SGLang for<br>OpenVINO]:::node-todo
     C2 --> D1
     C2 --> D1
    C3 --> D2
    %% C2 --> D3
    %% C3 --> D4
   end
   end


Line 116: Line 111:
   subgraph Models
   subgraph Models
     M1(Qwen3.8-27B)
     M1(Qwen3.8-27B)
    M2(Devstral-Small-2-24B)
     T2 --> M1
     T2 --> M1 & M2
   end
   end


Line 124: Line 118:
     V2(Optimize<br>lmsw.yaml)
     V2(Optimize<br>lmsw.yaml)
     C2 --> V1
     C2 --> V1
    C3 --> V1
     T2 --> V1
     T2 --> V1
     T4 ---> V1
     T4 ----> V1
     T6 ---> V2
     T6 ---> V2
   end
   end

Latest revision as of 03:12, 2 October 2026

Topics

🤖 Agent  🧠 Model  🚀 Engine  ⚙️ GPU & Driver  📊 Evaluation
ZooCode
axe
MCP
hf (Model Management)
Qwen Image 2.1
Bonsai 2 27B
Qwen 3.8 27B
Gemma 4
Devstral Small 2 24B
Experience
vLLM
llama-swap
ComfyUI
llama.cpp for SYCL
llama.cpp for CUDA
llama.cpp for OpenVINO
llama.cpp for TurboQuant
SGLang
Intel Arc Pro B70
Install driver
Install oneAPI (SYCL)
Install OpenVINO
.bashrc for B70
Experiment
NVIDIA RTX 3060
Install driver
AMD Ryzen AI 7 350
Experiment
Tools
Install nvtop
PCIe Trouble Shooting
Comparison
LM/GPU Comparison
Performance

News

Models management

  • llama-swap can launch model, cannot manage users.
  • LiteLLM cannot launch model, can manage users.

Investigating llama.cpp on Intel Arc Pro B70